IBM Granite under the hood
Docket fine-tunes IBM Granite’s 3B and 8B language models for this task. Granite supplies the language understanding; Docket adapts it for typed decisions and serves the results through a System One-compatible API.
IBM GRANITE · JEV-STYLE DECISIONS
Language in.
Decisions out.
Docket is a Jev-style decision model built on IBM Granite. Self-host 3B or 8B to classify inputs, route work, score relevance, and choose an agent’s next action—with typed answers and calibrated probabilities.
Jev-style models use language understanding to make bounded decisions. You provide context, a question, and the possible answers. The model scores those answers, giving your software a result it can use directly.
Docket fine-tunes IBM Granite’s 3B and 8B language models for this task. Granite supplies the language understanding; Docket adapts it for typed decisions and serves the results through a System One-compatible API.
Choice: select a tool, route, or label. Yes/no: estimate whether a condition holds. Score: rate an item on your ordered scale. Each uses probabilities over the valid answers, with per-type calibration.
Use the answers to branch application logic, prioritise items, or control an agent loop. One forward pass per question produces the answer distribution, without generating an explanation first.
Median decision latency
Median decision latency
1 × H100 SXM 80 GB
~1,000-token prompts · FP8
One request in flight · same-host HTTP
Use the same decision interface across application features, data pipelines, evaluation systems, and agent loops. Define the question and the outputs your software can use.
Assign a label, select a queue, or choose which workflow runs next.
Score relevance or quality, and prioritise items using your own scale.
Assess whether an output meets your criteria, or flag an item for review.
Choose the next tool, check an action, or decide whether to keep working.
These are intended workloads. Evaluate against your own labelled decisions before automating actions.
Docket scores your allowed answers in one forward pass per question. Choice, yes/no, and score replace a text-generation-and-parsing loop. You get probabilities for the alternatives, as well as the selected answer.
Provide context, instructions, and answer options for each question. The same API handles categorical choices, yes/no checks, and scores across different workloads.
Host the model and decision endpoint yourself. The model and server are Apache-2.0, and there is no hosted inference service required by Docket.
The request and response shapes follow TypeSafe System One. Existing clients can target a self-hosted Docket endpoint by changing their base URL.
3B and 8B scored higher on the sealed enterprise benchmark. Jev scored higher on JevBench and on the held-out workflow. These results describe the evaluated workloads; test the model on the decisions you plan to run.
Typical generative fine-tuning teaches a model to reproduce a response token by token. Docket trains at the decision point: given the context and question, put more probability on the correct allowed answer.
The training loss compares only valid answer tokens at the same position used by the serving engine. It targets the choice, yes/no answer, or score level directly, rather than the wording of a generated explanation.
LoRA adapts the Granite weights, and the adapters are merged for deployment. After training, separate temperatures for choice, yes/no, and score are fitted on held-out validation data to calibrate the probabilities. Calibration is a separate step from learning the decisions.
LOWER COST, LOWER LATENCY
A smaller decision model for workflows where latency and serving cost matter most.
HIGHER MEASURED ACCURACY
Higher accuracy on the reported decision benchmarks, with more compute per request.
Latency: ~1,000-token prompts, one H100 SXM 80 GB, FP8, one request in flight. Cost estimate: one L40S 48 GB, FP8, four requests in flight, assumed $2.61/hour at 70% average utilisation.
Both 3B and 8B use an 8,192-token context by default in the GPU server. That budget includes your context, question, instructions, answer options, and chat formatting. The vLLM backend reserves one answer token, leaving up to 8,191 tokens for the complete rendered prompt.
Longer-document evaluations use the 8K configuration. The accuracy figures displayed here were measured up to about 3,000 tokens; headline latency uses about 1,000 tokens.
3B and 8B · based on IBM Granite · Apache-2.0
Accuracy, generalisation, and serving performance answer different questions. Here’s the comparison for the two models.
| Benchmark | Docket 3B | Docket 8B | Jev reference |
|---|---|---|---|
| Enterprise decisions3,063 questions · weighted accuracy | 89.9% | 92.3% | 86.4% |
| JevBench231 public items · correct answers | 170 / 231 | 187 / 231 | 198 / 231 |
| Unseen workflowDecision type excluded from training | 60.7% | 64.0% | 70.8% |
Enterprise results use grouped splits and a sealed test opened after model selection. JevBench items were not used in training or selection. Accuracy was measured on documents up to about 3,000 tokens. Reference evaluation: October 3, 2026.
All benchmarks, methodology, and limitationsThe model is only part of the latency budget. The serving stack reduces compute, avoids repeated work, and keeps questions moving through the GPU.
Lower-precision inference increased L40S throughput by roughly 31–32% over bf16 at eight concurrent requests. FP8 and bf16 accuracy matched in the checks performed.
The server captures whole prefills up to 2,048 tokens by default. For a 250-token 3B prompt in the bf16 control, median engine time fell from 13.7 to 8.7 ms with the larger graph capture.
vLLM batches questions across concurrent requests. On the H100, four requests in flight reached 92.9 requests/s for 3B and 48.5 for 8B. Higher throughput comes with higher per-request latency.
Several questions about the same state can reuse cached prefix computation. With prefix caching enabled, three questions took 27.3 ms for 3B and 37.1 ms for 8B in the in-process study.
The engine reads probabilities only over valid answer tokens. There is no multi-token explanation to generate, and no generated JSON to repair.
Oversized inputs return HTTP 422 instead of silently dropping evidence. The server exposes health counters and a latency header, with a portable Transformers backend available for checks.
| Model | HTTP p50 | HTTP p95 | HTTP p99 | Requests/s |
|---|---|---|---|---|
| Docket 3B | 17.8 ms | 18.5 ms | 20.3 ms | 55.9 |
| Docket 8B | 27.2 ms | 28.2 ms | 28.9 ms | 36.8 |
GPU1 × NVIDIA H100 SXM 80 GB (Hopper)
InputOne choice question; ~1,000-token prompts. Workload mean: 987.4 tokens.
ConfigurationFP8; CUDA graph capture to 2,048 tokens; one request in flight.
BoundaryEnd-to-end HTTP on the same host, including server handling. External network travel is excluded. Unique requests; cached-prefix speedups are excluded from the headline results.
For comparison, the L40S 48 GB study measured 43 ms for 3B and 87 ms for 8B at one request in flight, with roughly 1,000-token prompts and FP8. These are measured GPU configurations, not guaranteed latency on another GPU or deployment.
Start with labelled examples from your own workflow. Set thresholds and human-review paths using those results.