Request and response format
The request and response format follows the TypeSafe System One API, so System One clients
work against docket-serve by changing the base URL.
Request
POST /v1/systemone (also served at /api/alpha/decisions)
{
"model": "optional, ignored",
"state": "any text, or a JSON object/array (rendered as indented JSON)",
"questions": {
"<question id>": {"type": "choice", "instructions": "...", "criteria": {"<key>": "<one-line description>", "...": "..."}},
"<question id>": {"type": "noul", "instructions": "...", "criteria": {"true": "optional: when to answer yes", "false": "optional"}},
"<question id>": {"type": "score", "instructions": "...", "criteria": ["level 0 description", "level 1 description", "..."]}
}
}
Limits:
- choice: 2 to 255 options. Descriptions are optional but improve accuracy.
- score: 2 to 10 levels, ordered lowest first.
- Questions: at least one per request. All questions about one state are scored
concurrently and share its cached prefix, so extra questions are cheap.
- Length (vLLM): both 3B and 8B default to an 8,192-token total context via
--max-model-len 8192. One token is reserved for the answer, leaving at most
8,191 prompt tokens per question. Count the complete rendered prompt: state
(text or structured data), question, instructions, options, and Granite chat formatting.
Longer requests get HTTP 422; nothing is truncated silently.
The portable Transformers (hf) backend allows 8,192 prompt tokens; this backend
currently uses its own fixed prompt limit rather than --max-model-len.
Longer-document evaluations use the 8K configuration. The accuracy figures on
this site were measured up to about 3,000 tokens; headline latency uses about 1,000.
Response
{
"model": "docket-3b",
"answers": {
"<id>": {"type": "choice", "choice": "<key>", "confidence": 0.0, "probabilities": {"<key>": 0.0}},
"<id>": {"type": "noul", "noul": 0.0},
"<id>": {"type": "score", "score": 1.7, "confidence": 0.0, "legend": {"0": "..."}, "probabilities": {"0": 0.0}}
},
"usage": {"input_tokens": 512, "output_tokens": 3}
}
| Field | Meaning |
|---|---|
noul |
P(yes) |
choice |
the most probable key |
confidence |
for choice: (p_max - 1/K) / (1 - 1/K) (0 = uniform, 1 = certain); for score: 1 - std / std_max |
score |
the probability-weighted level index |
probabilities |
calibrated with the per-type temperatures in docket_calibration.json (start the server with --no-calibration for raw probabilities) |
usage.input_tokens |
prompt tokens processed |
usage.output_tokens |
number of questions answered |
Errors. Validation and length errors return HTTP 422 with
{"detail": [{"loc": ["body"], "msg": "...", "type": "value_error"}]}.
Other endpoints
| Endpoint | Returns |
|---|---|
GET /v1/models |
the served model name |
GET /healthz, GET /health |
{"ok": true, "model": ..., "requests": n, "questions": n, "input_tokens": n} |
The response header x-docket-latency-ms gives the server-side time per request.
Server options
docket-serve --model PATH_OR_HF_ID [--backend vllm|hf] [--port 8080] [--quantization fp8]
[--max-model-len 8192] [--max-num-seqs 256] [--gpu-memory-utilization 0.9]
[--cudagraph-max-tokens 2048] [--name docket-3b] [--no-calibration]
Environment equivalents: DOCKET_MODEL, DOCKET_NAME, DOCKET_BACKEND, DOCKET_QUANTIZATION, DOCKET_CUDAGRAPH_MAX_TOKENS, PORT.
Latency. For the lowest latency use an H100-class GPU with --quantization fp8, and keep
1-2 requests in flight per replica. Measured p50 at ~1,000 tokens: Docket-3B 17.8 ms,
Docket-8B 27.2 ms. --cudagraph-max-tokens (default 2048) runs whole prefills up to that
length as CUDA graphs, which cuts fixed overhead on short prompts.
Throughput. For about 1,000-token decisions, 4-8 requests in flight per GPU reach peak throughput; more only adds latency. See the model cards for measured numbers.