PROJECT 001 · OPEN SOURCE · APACHE-2.0

Docket.

Fast models for decision workloads.
3B and 8B models you can host yourself.

Benchmark conditions

These measurements cover Docket 3B and 8B. Decision accuracy and serving speed are separate studies. The accuracy tables use the same labelled items and scoring method for each compared model. Serving measurements identify the GPU, input length, precision, concurrency, and timing boundary.

Quality reference evaluation: October 3, 2026. H100 serving study: October 4, 2026.

Enterprise decisions

A sealed test of 3,063 questions across 10 enterprise workflows. The metric is workflow-weighted exact-label accuracy. Splits were grouped, and the test was opened after model selection. Accuracy was measured on documents up to about 3,000 tokens.

Model Base model accuracy Docket accuracy Difference vs. Jev
Docket 3B 55.6% 89.9% +3.4 points (95% CI +1.7 to +5.0)
Docket 8B 63.1% 92.3% +5.9 points (95% CI +4.0 to +7.7)
Jev reference — 86.4% Reference

The difference intervals use paired group bootstrap. These results describe the evaluated enterprise workflows; they do not establish accuracy on a new business process.

JevBench

231 public items, with the same harness for every row. No JevBench item was used in training or model selection.

Model Easy / 48 Original / 72 Hard / 111 Total / 231
Granite 3B base 48 51 54 153
Granite 8B base 48 64 49 161
Docket 3B 48 66 56 170
Docket 8B 48 71 68 187
Jev reference 48 70 80 198

Generalisation and confidence

The held-out workflow uses a decision type excluded from training. The high-confidence error metric is the share of all labelled answers that were wrong while reported confidence was at least 0.9; it is not the error rate among only high-confidence answers.

Model Unseen workflow accuracy Wrong at ≥0.9 confidence, enterprise test
Docket 3B 60.7% 1.8%
Docket 8B 64.0% 1.8%
Jev reference 70.8% 1.1%

Docket can be overconfident on unfamiliar decisions. On the unseen workflow, roughly a quarter of Docket’s labelled answers were wrong despite confidence of at least 0.9. The calibrated probabilities are useful signals, not proof that a decision is correct. Choose production thresholds using independently labelled examples from your own workflow.

HTTP latency

Hardware: one NVIDIA H100 SXM 80 GB (Hopper). Workload: one choice question over approximately 1,000 input tokens; the workload mean was 987.4 tokens. Serving: FP8, CUDA graph capture up to 2,048 tokens.

Timing is end-to-end HTTP with the client on the same host. It includes server handling but excludes external network travel. Requests are unique; cached-prefix speedups are excluded from the headline results. There were zero recorded request errors in these runs.

Model Requests in flight Mean input tokens p50 p95 p99 Requests/s
Docket 3B 1 987.2 17.8 ms 18.5 ms 20.3 ms 55.9
Docket 3B 2 987.0 26.2 ms 27.4 ms 32.8 ms 75.8
Docket 3B 4 987.5 43.1 ms 43.7 ms 44.8 ms 92.9
Docket 8B 1 987.3 27.2 ms 28.2 ms 28.9 ms 36.8
Docket 8B 2 987.3 45.1 ms 46.0 ms 46.7 ms 44.4
Docket 8B 4 987.3 82.5 ms 84.0 ms 85.2 ms 48.5

One request in flight delivers the lowest measured latency. Increasing concurrency raises throughput and adds queueing and batching latency. These short benchmark runs do not establish a production SLA.

Latency by prompt length

Same H100 SXM 80 GB, FP8, and 2,048-token graph capture. Prefix caching is disabled. These are engine-only p50 measurements; rendering, tokenisation, HTTP, and network time are excluded. Each length/model cell contains 50 measurements. The nominal 1,000-token condition averaged 988.1 tokens.

Target prompt tokens Measured mean tokens 3B engine p50 8B engine p50
250 249.9 8.8 ms 11.2 ms
500 499.9 9.8 ms 14.1 ms
1,000 988.1 14.6 ms 23.3 ms
2,000 2000.9 23.9 ms 42.5 ms

Serving throughput

Hardware: one NVIDIA L40S 48 GB, with the serving process limited to four vCPUs. Workload: approximately 1,000 input tokens per decision. Precision: FP8.

Costs are a study estimate using a GPU rate of $2.61/hour at 70% average utilisation, not a current provider quote. The denominator is one million input tokens.

Model Requests in flight Requests/s HTTP p50 Estimated $ / M input tokens
Docket 3B 1 23.4 43 ms $0.045
Docket 3B 4 32.6 123 ms $0.032
Docket 3B 8 35.1 247 ms $0.030
Docket 8B 1 11.5 87 ms $0.091
Docket 8B 4 15.1 265 ms $0.069
Docket 8B 8 15.7 548 ms $0.067

At one request in flight, the H100 measured 17.8 ms versus the L40S’s 43 ms for 3B, and 27.2 ms versus 87 ms for 8B. Hardware and serving configuration are part of the result; another GPU is not assumed to produce the same latency.

FP8 and CUDA graph improvements

At eight requests in flight on the L40S, FP8 improved throughput relative to bf16:

Model bf16 requests/s FP8 requests/s Improvement
Docket 3B 26.8 35.1 ~31%
Docket 8B 11.9 15.7 ~32%

FP8 and bf16 accuracy matched in the checks performed. This is not a claim that every future workload or quantisation configuration will have identical accuracy.

On the H100, enlarging CUDA graph capture from the default short-prefill range to 2,048 tokens reduced the 250-token 3B bf16 engine median from 13.7 ms to 8.7 ms. For approximately 1,000 tokens, the optimized FP8 engine medians were 14.6 ms for 3B and 23.3 ms for 8B. Full-plus-piecewise graph mode added no measurable improvement beyond the larger capture size.

Several questions, one state

The original H100 FP8 in-process run scored three questions concurrently and reused their shared state prefix. Prefix caching was enabled. These numbers include preparation and engine work, but exclude HTTP and external network time. The three-question condition contains 35 measurements per model.

Model One question p50 Three questions, same state p50
Docket 3B 17.9 ms 27.3 ms
Docket 8B 27.2 ms 37.1 ms

This is a prefix-reuse scenario, not a cold independent-request latency claim. Keep reusable instructions or shared context at the start of the state when your workload permits it.

Deployment boundaries

  • Your workflow: validate on your own labelled inputs, then set confidence thresholds and review paths.

  • Input length: both model sizes default to an 8,192-token total serving context in vLLM, with one answer token reserved (up to 8,191 tokens for the complete rendered prompt). The context, question, instructions, options, and chat formatting all count. Longer-document evaluations use the 8K configuration; the accuracy figures reported on this page were evaluated up to about 3,000 tokens. Headline latency was measured around 1,000 tokens.

  • Language: the reported decision-quality evaluations cover English.

  • Untrusted text: Docket is not a security boundary. Do not treat a model decision as a substitute for deterministic access controls.

  • Timing: local engine, local HTTP, and remote end-to-end latency are different measurements. The tables label the boundary for each study.

See the API reference for request shapes, limits, and serving options.