Builder mode
API endpoints
Call your agent from your own code or any OpenAI-compatible app, using a key of your own.
SettingsAPI gives your agent a web address that speaks the OpenAI chat format, and a key to go with it. Your own code, or any app that can talk to an OpenAI-compatible service, can then send your agent messages and get its replies back.
Getting your address and key
- Open SettingsAPI (open it). In builder mode it's under "Developer" in the Settings list.
- The panel shows "Base URL" with a "Copy" button. Until you've made a key, a line under it says there's no key yet.
- Press "Generate key". It reads "Working…" for a moment, then an "API key" field appears with its own "Copy" button.
- The Base URL is the address you opened the dashboard on, followed by
/api/instance/openai/v1. Copy it from the panel rather than typing it. - The key starts with
mh_pk_, followed by 48 letters and digits. - The key field is hidden like a password. "Copy" puts the key on your clipboard and reads "Copied" for a moment. The key stays in the panel, so you can come back and copy it again.
- A ready-made setup appears under "cursor / openai client" once you have a key: your Base URL, your key, and the model name
hermes-agent.
If the panel can't load, it says "Couldn't load your API endpoint settings." If making or replacing a key fails, it says "Couldn't update the key. Try again."
Rotating the key
Press "Rotate key". The old key stops working straight away and a new one takes its place. Put the new key into every app that used the old one.
- This key is only for the API. Replacing it doesn't affect the web chat or Telegram.
- There's no switch to turn the endpoint off. Rotating is how you cut off a key that got out; then don't share the new one.
Sending a message with curl
Put your Base URL and key into the first two lines, then run this in a terminal:
BASE_URL='<your Base URL>'
KEY='<your API key>'
curl -sS "$BASE_URL/chat/completions" \
-H "Authorization: Bearer $KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"hermes-agent","messages":[{"role":"user","content":"Say hello in one sentence."}]}'
The reply comes back as OpenAI-style JSON. Hermes's answer is the text at choices[0].message.content.
To get the answer as it's written, add "stream": true to the body. The reply then arrives as a stream of server-sent events; add -N to the curl command to see it as it comes in.
Using an OpenAI SDK
With the official OpenAI library for Python, set the base URL and key on the client:
from openai import OpenAI
client = OpenAI(
base_url="<your Base URL>",
api_key="<your API key>",
max_retries=0,
)
reply = client.chat.completions.create(
model="hermes-agent",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(reply.choices[0].message.content)
The Base URL already ends in /v1, which is what the library expects, so paste it as it is.
max_retries=0 turns off the library's automatic retries. By default it sends a request up to two more times after a 429, a server error or a timeout, and a retry can run your agent again. A request cut off at 300 seconds could run it three times.
In other apps that let you set an OpenAI base URL, enter the same three things the panel's sample shows: the Base URL, your key as the API key, and hermes-agent as the model.
What the endpoint accepts
- Addresses under the Base URL only.
GETandPOSTrequests to paths under/v1are passed on to your agent. A request with a key to any other path gets404with{"error":"not_found"}; with no key it gets401. OnlyGETandPOSTare accepted. - The key as a bearer token: an
Authorization: Bearer mh_pk_…header on every request. No dashboard login is needed. - Two headers go through to your agent,
Content-TypeandAccept. Others are dropped. - A request runs your agent the same way a message in the web chat does, with its skills, tools and connectors.
The endpoint doesn't send a chat session along with a request. Don't count on a conversation through the API showing up in your web chat, or on it picking up where a chat left off. Put everything a request needs in its messages.
Limits
- Your plan's running-time cap applies, just as in the web chat. Once it's reached, requests get
429until the window frees up. Plans and credits explains the cap. - A sleeping agent is woken first. If it isn't ready within 180 seconds, the request gets
503withinstance_warming_up; send it again. - Each request can run for up to 300 seconds, counting the time it takes to wake your agent. A longer request is cut off.
- After your plan has ended, requests get
403. Once your agent has been removed, the key no longer works at all and requests get401.
Responses and errors
A successful request returns whatever your agent sends back, with its own status code. So does an error from your agent itself. These come from MyHermes before a request reaches your agent:
| Status | Body | What it means |
|---|---|---|
401 | {"error":"unauthorized"} | No key, a wrong key, or a key that's been rotated. |
403 | {"error":"subscription_cancelled"} (or no_subscription) | Your plan has ended. |
404 | {"error":"not_found"} | The path isn't under the Base URL's /v1. |
429 | {"error":"quota_exceeded","reason":…,"resetAt":…} | The running-time cap is reached. reason is 5h_exceeded or 7d_exceeded; resetAt is when it frees up, as a Unix timestamp in milliseconds. |
503 | {"error":"instance_warming_up","detail":…} | Your agent didn't wake within 180 seconds. |
503 | {"error":"quota_unavailable"} or {"error":"subscription_unavailable"} | MyHermes couldn't check your plan just then. Try again. |
502 | {"error":"upstream_unreachable"} or {"error":"resolve_failed"} | MyHermes couldn't reach your agent, or couldn't look up your key. Try again. |
Keeping the key safe
- The key is all it takes. Anyone who has it can send your agent requests, and those requests run with your agent's tools and connectors and count toward your running-time cap.
- Keep it out of shared code, public repositories and screenshots.
- If it gets out, press "Rotate key" straight away.
If something goes wrong
401. Check you copied the whole key, starting withmh_pk_. If you or someone else pressed "Rotate key", copy the new key from the panel.404withnot_found. Check the address starts with your Base URL, for example<your Base URL>/chat/completions.503withinstance_warming_up. Your agent was asleep and took too long to wake. Send the request again.429withquota_exceeded. Wait until the time inresetAt, or see Plans and credits.- The request stops after about five minutes. Requests are cut off at 300 seconds. Ask for less in each request.
Common questions
Is this the same agent that answers me in the chat and on Telegram?
Yes. Requests reach your own agent, the same one behind the web chat and your Telegram bot. Just don't count on sharing a conversation with the chat: see What the endpoint accepts.
Can I turn the endpoint off?
There's no off switch. Press "Rotate key" to cut off the current key, and keep the new one to yourself.
Which model name should I send?
Use hermes-agent, as the panel's sample shows.
What's next
- Webhooks: let other services wake your agent, with replies in Telegram.
- Plans and credits: the running-time cap and what each plan includes.
- Connectors: apps your agent can use when it handles a request.