api reference
One endpoint. OpenAI-compatible. Cheap.
If your code speaks to the OpenAI API, it already speaks to ours. Point your client at our base URL and pick a model.
01Base URL
https://hitujvywvh.us-east-2.awsapprunner.com/v1this is the live production endpoint. Keep the /v1 suffix when setting your OpenAI-compatible client's base URL.
02Authentication
Send your API key as a standard bearer token on every request:
Authorization: Bearer $GOODENOUGH_API_KEY03Models
gpt-oss-120b is the flagship — start there. It's a reasoning model: it spends tokens thinking (returned as reasoning_content) before the final answer, so give it generous max_tokens (500+) or the answer may arrive truncated or empty.
Prices are USD per million tokens. Send model: "auto" and the router picks the right model per request — the response's model field tells you what it chose.
Auto uses GPT-OSS for general chat, Qwen3 Coder Next for code, and Qwen3 VL for images. Kimi K2.5, GLM 5, and Kimi K3 are available for explicit model selection.
| model id | name | source | input /M | output /M |
|---|---|---|---|---|
| auto | Auto (router) | — | $0.23 | $0.92 |
| gpt-oss-120b | GPT-OSS 120B | na | $0.23 | $0.92 |
| qwen3-coder-next | Qwen3 Coder Next | cn | $0.77 | $1.85 |
| qwen3-vl | Qwen3 VL 235B-A22B | cn | $0.82 | $4.09 |
| kimi-k3 | Kimi K3 (Flagship) | cn | $5.00 | $25.00 |
| kimi-k2.5-bedrock | Kimi K2.5 | cn | $0.92 | $4.62 |
| glm-5-bedrock | GLM 5 | cn | $1.54 | $4.92 |
| typical frontier API (for reference) | $3.00 | $15.00 | ||
Auto routing (IRL)
The router is a learned classifier (embedding + linear head) with deterministic constraint filters. With model: "auto" the request body may carry an optional preferences.exclude_sources field — a hard filter on model origin (values from the source column above) applied before the router picks:
curl https://hitujvywvh.us-east-2.awsapprunner.com/v1/chat/completions \
-H "Authorization: Bearer $GOODENOUGH_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "refactor this function"}],
"preferences": {"exclude_sources": ["cn"]}}'The response's model field reveals which model the router picked — no custom headers. If the filter excludes every servable model, the request fails with a 400.
Attachments & vision
Messages may carry OpenAI-style content parts: an array mixing text and image_url parts (data: URLs work). With model: "auto", any request carrying an image routes to the vision pool (qwen3-vl) — text pools never see images, and the vision pool only serves image requests:
curl https://hitujvywvh.us-east-2.awsapprunner.com/v1/chat/completions \
-H "Authorization: Bearer $GOODENOUGH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBOR..."}}
]}]
}'The vision pool is optional per deployment; when it isn't deployed (or the provenance filter excludes it), image requests fail with a 400 naming the constraint.
Tool calling
Standard OpenAI function calling — tools + tool_choice (including "auto") — works on the production text models, streaming included, and passes through model: "auto" routing untouched:
resp = client.chat.completions.create(
model="gpt-oss-120b", # best tool-use judgment; see notes below
messages=[{"role": "user", "content": "What's the weather in Toronto?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}],
tool_choice="auto",
)
call = resp.choices[0].message.tool_calls[0]
# call.function.name == "get_weather"
# call.function.arguments == '{"city": "Toronto"}'Model notes, from our own test suite: gpt-oss-120b is the recommendation for agentic workloads — it calls when it should, declines when it shouldn't, and handles tool-result turns and streaming cleanly. Qwen3 Coder Next handles code-focused requests; Qwen3 VL is reserved for image-bearing requests. Tool behavior can vary by provider model, so applications should validate tool arguments before execution.
04curl
curl https://hitujvywvh.us-east-2.awsapprunner.com/v1/chat/completions \
-H "Authorization: Bearer $GOODENOUGH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "write a function that dedupes a list"}],
"stream": true,
"stream_options": {"include_usage": true}
}'05Python (openai client)
from openai import OpenAI
client = OpenAI(
base_url="https://hitujvywvh.us-east-2.awsapprunner.com/v1",
api_key="YOUR_GOODENOUGH_KEY",
)
resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize this support ticket: ..."}],
)
print(resp.model) # what the router picked, e.g. "gpt-oss-120b"
print(resp.choices[0].message.content)06JavaScript — change your baseURL
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://hitujvywvh.us-east-2.awsapprunner.com/v1", apiKey: process.env.GOODENOUGH_API_KEY });
// done — every existing chat.completions call now runs on Good Enough07Errors
Errors come back as OpenAI-style JSON: an error object with message, type and code. The three you are most likely to meet:
401 — invalid or missing key
The Authorization header is missing, malformed, or carries a key we don't recognize (including revoked keys). Send the bearer header exactly as in section 02. Keys are shown once at creation — if yours is lost, mint a new one in the console.
400 — out of credits (budget_exceeded)
When your account's credit balance is used up, every request is refused with error.type: "budget_exceeded":
{
"error": {
"message": "ExceededBudget: User=you@example.com over budget. Spend=5.0041, Budget=5.0",
"type": "budget_exceeded",
"code": "400"
}
}top up in the console; requests resume immediately.
429 — rate limited
Keys default to 60 requests/min (spend is separately bounded by your credit balance). On a 429, wait a moment and retry — exponential backoff starting around one second clears it. If you need a higher sustained rate, get in touch.
Streaming works the standard way: set stream: true (and stream_options.include_usage to get a final usage chunk). SSE lines arrive as data: JSON, terminated by data: [DONE].