LLMEndpointProbeResult
Structured result of a GET {base_url}/models probe.
reachable AND authenticated (2xx); also requires the inference step to pass when it was requested
ok | auth_failed | http_error | unreachable | invalid_url | inference_failed
Human-readable connectivity result. Clients may localize it using the structured reason value.
status_code object
upstream HTTP status when a response was received
- integer
- null
probed_url object
model-list URL actually requested — lets the caller confirm what was probed without replicating per-provider path assembly
- string
- null
probed_provider object
provider actually probed
- string
- null
sorted names of the headers this probe carries — derived auth plus custom, after custom overrides. Names only; the values are write-only credentials and are never exposed. Empty when no api_key resolved and no custom headers applied
true when the stored model's base_url/provider/credential were probed instead of the ones in the request body (stored-key branch) — the caller's base_url was NOT tested
falseinference object
Result of the generation round trip, or null when inference verification was not requested or the model-list probe failed before it could run.
- LLMInferenceProbeResult
- null
the completion request returned 2xx
ok | auth_failed | http_error | unreachable | invalid_url | model_required
Human-readable inference result in English. Clients may localize it using the structured reason value.
status_code object
upstream HTTP status when a response was received
- integer
- null
probed_url object
completion URL actually requested
- string
- null
body object
Upstream response body, masked and length-capped. It is also returned when the request fails, and is null when no response was received.
- string
- null
content object
generated text, masked and length-capped
- string
- null
reasoning_content object
chain of thought when the endpoint reports one (Anthropic thinking blocks, OpenAI-compatible reasoning_content/reasoning)
- string
- null
usage object
token usage block the endpoint reported, as-is
- object
- null
finish_reason object
OpenAI finish_reason / Anthropic stop_reason, when reported
- string
- null
ttft_ms object
Time to the first token in milliseconds, measured from request start to the first streamed text or thinking delta. It is null unless measure_ttft was requested, and also when the stream carried no delta.
- number
- null
models object
model ids returned by the endpoint, when parseable
- string[]
- null
model_found object
Whether the requested model appears in the returned model list, or null when no model was checked.
- boolean
- null
limits object
Token limits reported for the requested model, or null when the endpoint reports no limits or the requested model is unavailable.
- LLMDetectedLimits
- null
context_window object
detected context window (tokens)
- integer
- null
max_tokens object
detected max output tokens cap
- integer
- null
reasoning_levels object
Supported reasoning-effort levels advertised by the endpoint, or null when the endpoint reports no reasoning capability.
- string[]
- null
Possible values: [none, minimal, low, medium, high, xhigh, max]
{
"success": true,
"reason": "string",
"message": "string",
"latency_ms": 0,
"status_code": 0,
"probed_url": "string",
"probed_provider": "string",
"probed_header_names": [
"string"
],
"used_stored_endpoint": false,
"inference": {
"success": true,
"reason": "string",
"message": "string",
"latency_ms": 0,
"status_code": 0,
"probed_url": "string",
"body": "string",
"content": "string",
"reasoning_content": "string",
"usage": {},
"finish_reason": "string",
"ttft_ms": 0
},
"models": [
"string"
],
"model_found": true,
"limits": {
"context_window": 0,
"max_tokens": 0,
"reasoning_levels": [
"none"
]
}
}