IBM GRANITE · JEV-STYLE DECISIONS

Docket.

Language in.
Decisions out.

Docket is a Jev-style decision model built on IBM Granite. Self-host 3B or 8B to classify inputs, route work, score relevance, and choose an agent’s next action—with typed answers and calibrated probabilities.

01 / MEET DOCKET

A language model.
A decision interface.

Jev-style models use language understanding to make bounded decisions. You provide context, a question, and the possible answers. The model scores those answers, giving your software a result it can use directly.

IBM Granite under the hood

Docket fine-tunes IBM Granite’s 3B and 8B language models for this task. Granite supplies the language understanding; Docket adapts it for typed decisions and serves the results through a System One-compatible API.

Three ways to ask

Choice: select a tool, route, or label. Yes/no: estimate whether a condition holds. Score: rate an item on your ordered scale. Each uses probabilities over the valid answers, with per-type calibration.

A result your workflow can act on

Use the answers to branch application logic, prioritise items, or control an agent loop. One forward pass per question produces the answer distribution, without generating an explanation first.

DOCKET 3B17.8 ms

Median decision latency

DOCKET 8B27.2 ms

Median decision latency

MEASURED ON

1 × H100 SXM 80 GB
~1,000-token prompts · FP8
One request in flight · same-host HTTP

Full test conditions
02 / DECISION WORKLOADS

Make decisions
across your stack.

Use the same decision interface across application features, data pipelines, evaluation systems, and agent loops. Define the question and the outputs your software can use.

01

Classification & routing

Assign a label, select a queue, or choose which workflow runs next.

02

Ranking & scoring

Score relevance or quality, and prioritise items using your own scale.

03

Evaluation & checks

Assess whether an output meets your criteria, or flag an item for review.

04

Agent orchestration

Choose the next tool, check an action, or decide whether to keep working.

These are intended workloads. Evaluate against your own labelled decisions before automating actions.

03 / WHAT MAKES IT DIFFERENT

Designed around
the answer you need.

A typed answer, directly

Docket scores your allowed answers in one forward pass per question. Choice, yes/no, and score replace a text-generation-and-parsing loop. You get probabilities for the alternatives, as well as the selected answer.

Your task, defined at request time

Provide context, instructions, and answer options for each question. The same API handles categorical choices, yes/no checks, and scores across different workloads.

Your infrastructure, your data path

Host the model and decision endpoint yourself. The model and server are Apache-2.0, and there is no hosted inference service required by Docket.

A familiar decision API

The request and response shapes follow TypeSafe System One. Existing clients can target a self-hosted Docket endpoint by changing their base URL.

Compared with the hosted Jev reference in our tests

3B and 8B scored higher on the sealed enterprise benchmark. Jev scored higher on JevBench and on the held-out workflow. These results describe the evaluated workloads; test the model on the decisions you plan to run.

04 / HOW IT LEARNS

Fine-tuned for
the decision itself.

Typical generative fine-tuning teaches a model to reproduce a response token by token. Docket trains at the decision point: given the context and question, put more probability on the correct allowed answer.

Train what you serve

The training loss compares only valid answer tokens at the same position used by the serving engine. It targets the choice, yes/no answer, or score level directly, rather than the wording of a generated explanation.

Adapt Granite, then calibrate

LoRA adapts the Granite weights, and the adapters are merged for deployment. After training, separate temperatures for choice, yes/no, and score are fitted on held-out validation data to calibrate the probabilities. Calibration is a separate step from learning the decisions.

05 / TWO MODEL SIZES

Choose the size
that fits the job.

LOWER COST, LOWER LATENCY

Docket 3B

A smaller decision model for workflows where latency and serving cost matter most.

Serving context
8,192 tokens
Enterprise accuracy
89.9%
H100 HTTP p50
17.8 ms
L40S serving estimate
$0.032 / M tokens

HIGHER MEASURED ACCURACY

Docket 8B

Higher accuracy on the reported decision benchmarks, with more compute per request.

Serving context
8,192 tokens
Enterprise accuracy
92.3%
H100 HTTP p50
27.2 ms
L40S serving estimate
$0.069 / M tokens

Latency: ~1,000-token prompts, one H100 SXM 80 GB, FP8, one request in flight. Cost estimate: one L40S 48 GB, FP8, four requests in flight, assumed $2.61/hour at 70% average utilisation.

8K context for both sizes

Both 3B and 8B use an 8,192-token context by default in the GPU server. That budget includes your context, question, instructions, answer options, and chat formatting. The vLLM backend reserves one answer token, leaving up to 8,191 tokens for the complete rendered prompt.

Longer-document evaluations use the 8K configuration. The accuracy figures displayed here were measured up to about 3,000 tokens; headline latency uses about 1,000 tokens.

Model weights on Hugging Face

3B and 8B · based on IBM Granite · Apache-2.0

Hugging Face Link coming soon
06 / THE EVIDENCE

Measured where
the decisions happen.

Accuracy, generalisation, and serving performance answer different questions. Here’s the comparison for the two models.

BenchmarkDocket 3BDocket 8BJev reference
Enterprise decisions3,063 questions · weighted accuracy89.9%92.3%86.4%
JevBench231 public items · correct answers170 / 231187 / 231198 / 231
Unseen workflowDecision type excluded from training60.7%64.0%70.8%

Enterprise results use grouped splits and a sealed test opened after model selection. JevBench items were not used in training or selection. Accuracy was measured on documents up to about 3,000 tokens. Reference evaluation: October 3, 2026.

All benchmarks, methodology, and limitations
07 / BUILT FOR SERVING

A shorter path
from input to decision.

The model is only part of the latency budget. The serving stack reduces compute, avoids repeated work, and keeps questions moving through the GPU.

01

FP8 inference

Lower-precision inference increased L40S throughput by roughly 31–32% over bf16 at eight concurrent requests. FP8 and bf16 accuracy matched in the checks performed.

02

CUDA graphs for longer prefills

The server captures whole prefills up to 2,048 tokens by default. For a 250-token 3B prompt in the bf16 control, median engine time fell from 13.7 to 8.7 ms with the larger graph capture.

03

Continuous batching

vLLM batches questions across concurrent requests. On the H100, four requests in flight reached 92.9 requests/s for 3B and 48.5 for 8B. Higher throughput comes with higher per-request latency.

04

Shared-prefix reuse

Several questions about the same state can reuse cached prefix computation. With prefix caching enabled, three questions took 27.3 ms for 3B and 37.1 ms for 8B in the in-process study.

05

Restricted answer readout

The engine reads probabilities only over valid answer tokens. There is no multi-token explanation to generate, and no generated JSON to repair.

06

Explicit serving limits

Oversized inputs return HTTP 422 instead of silently dropping evidence. The server exposes health counters and a latency header, with a portable Transformers backend available for checks.

Serving throughput and optimisation measurements
08 / LATENCY WITH CONTEXT

Know what
the numbers include.

ModelHTTP p50HTTP p95HTTP p99Requests/s
Docket 3B17.8 ms18.5 ms20.3 ms55.9
Docket 8B27.2 ms28.2 ms28.9 ms36.8

GPU1 × NVIDIA H100 SXM 80 GB (Hopper)

InputOne choice question; ~1,000-token prompts. Workload mean: 987.4 tokens.

ConfigurationFP8; CUDA graph capture to 2,048 tokens; one request in flight.

BoundaryEnd-to-end HTTP on the same host, including server handling. External network travel is excluded. Unique requests; cached-prefix speedups are excluded from the headline results.

For comparison, the L40S 48 GB study measured 43 ms for 3B and 87 ms for 8B at one request in flight, with roughly 1,000-token prompts and FP8. These are measured GPU configurations, not guaranteed latency on another GPU or deployment.

PUT IT TO WORK

Bring your workload.
Test your decisions.

Start with labelled examples from your own workflow. Set thresholds and human-review paths using those results.