PROJECT 001 · OPEN SOURCE · APACHE-2.0

Docket — Dataset.

Fast models for decision workloads.
3B and 8B models you can host yourself.

Data for decisions

Docket is fine-tuned on proprietary decision data and public language-understanding and mathematics datasets. The data connects language understanding to a specific decision: classify an input, check a condition, assess a proposed action, or select a level on a scale.

The categories below describe the data represented across Docket 3B and 8B. They cover enterprise workflows, reasoning over evidence, and general language tasks.

All examples on this page are illustrative. They explain the task types rather than reproduce training records. The answers show the intended labels, not recorded model predictions or benchmark results.

Three answer types

An example supplies context, a question, the allowed answers, and a correct label. The answer has one of three forms:

Answer type What it represents Illustrative example
Choice One category from the available options. A ticket says, “I was charged twice.” Choose billing, technical support, or sales. Label: billing.
Yes/no (noul) Whether a stated condition holds. The account has read-only access. “May it delete this file?” Label: no.
Score A level on an ordered scale. Severity is low, medium, or high. High means a complete service outage. All customer requests are failing. Label: high.

These answer types appear across different tasks. An incident can have a severity score, a yes/no escalation check, and a choice of response action.

Enterprise decisions

This part of the data represents operational decisions over documents, messages, records, and proposed actions. The question supplies the criteria that apply to the decision.

Task type Illustrative context and question Intended answer
Access and permissions Exporting a report requires export permission. The user has view permission only. May the user export it? No
Agent action checks An agent proposes deleting a file through a tool limited to read-only operations. May the action run? No
Document and agreement checks An agreement requires 30 days’ notice. A proposed change gives 10 days. Does the change satisfy that requirement? No
Expense checks The meal limit is $40 and the receipt is $32. Is the amount within the limit? Yes
Incident severity Critical means the whole service is unavailable. Every customer request is failing. Which severity applies? Critical
Claims and coverage Coverage includes device failures and excludes lost devices. The device was lost. Is the claim covered? No
Invoice matching An order is for four units at $25 each. The invoice totals $110. Does the amount match? No
Leave and balance checks The record shows three days of leave available. The request is for five days. Is the balance sufficient? No
Data-sharing checks A public destination accepts only redacted records. The proposed record contains an unmasked phone number. May it be shared there? No
Ticket routing A customer reports a duplicate charge. Choose the billing, technical support, or sales queue. Billing

The common structure is a decision grounded in the supplied context. Some questions ask for an action; others check one requirement or assign an ordered level.

Reasoning over evidence

These tasks require combining facts, following conditions, checking calculations, or deciding what the available evidence supports.

Task type Illustrative context and question Intended answer
Dates and deadlines A request arrives Monday and is due two business days later. There are no holidays. Which day is the deadline? Wednesday
Probability and expected cost Option A costs $100 plus a 20% chance of a $100 penalty. Option B costs $115 with no penalty. Which has the lower expected cost? Option B
Linked records An activity log names user7. The directory maps user7 to Alex in Payments; Morgan leads Payments. Who leads the actor’s team? Morgan
Conditions and exceptions Access requires identity verification and completed training. A suspension blocks access. Both requirements are met, but the account is suspended. May access proceed? No
Calculation verification An $80 bill has a 25% tip and is split equally among five people. A proposed answer says $20 each. Is that calculation correct? Yes
Logical conclusion checks · 3B Assignment requires completed training. Completion is not recorded. A claim says the person qualifies. Choose supported, contradicted, or insufficient evidence. Insufficient evidence

The logical conclusion-check category is represented in the 3B training data. It distinguishes a supported conclusion, a contradicted conclusion, and one that cannot be resolved from the supplied facts.

Public datasets used in fine-tuning

The training data also includes these public datasets. They represent reading comprehension, inference, intent classification, and mathematical reasoning. The links point to their dataset cards; the examples remain illustrations written for this page.

Dataset and task type Illustrative decision
BoolQ · passage-based yes/no questions A passage states that the library is closed on Sundays. “Is it open on Sunday?” Answer: no.
WANLI · natural-language inference Premise: all approved invoices have been paid. Claim: an approved invoice remains unpaid. Choose entailment, contradiction, or neutral. Answer: contradiction.
Banking77 · customer-query intent classification “I cannot find my bank card.” Select the matching customer-service intent. Answer: lost or stolen card.
GSM8K · arithmetic word problems Six boxes contain four pens each. Five pens are given away. Which numerical answer gives the number left? Answer: 19.
AQuA-RAT · mathematical multiple-choice problems Three identical subscriptions plus a $2 fee cost $14. Which option gives the price of one subscription? Answer: $4.

The context comes in different forms

The decision data includes prose such as emails and tickets, structured records and tables, and information containing amounts, dates, or explicit requirements. Questions can ask for a direct fact, combine several facts, or check a proposed conclusion. Some inputs leave relevant information unstated, so the task may require an insufficient-evidence answer rather than an assumption.

Dataset coverage describes the tasks present in fine-tuning. For measured performance, see the benchmarks. For the request and answer formats, see the API reference.