Myles decision models

One job, done well, on any machine.

Myles reads a situation, a question and the possible answers, and returns a probability for every option in a single pass. It classifies, routes, checks and approves: the decisions inside every AI product.

Myles sitting on a rock, checking his phone, beside a server rack
91.8%correct on 1,400 held-out decisions, ahead of the medium tier of a leading hosted LLM family
1/450thof the cost per decision, on a single CPU thread
0GPUs, cloud calls or data leaving your network

Myles-400M on our frozen evaluation set. Method below.

The family

ModelRuns onParametersTypical speed
Myles-22MYour browser or phone23M27 ms
Myles-150MA laptop149M196 ms
Myles-400MThe cheapest server you have395Mabout 500 ms
Myles-LOne 24 GB GPU14.7B97 ms
Myles-1BIn training1B—

Median time for one decision on one CPU thread of a desktop Intel processor: 22M and 150M as 8-bit builds, both timed while the machine was busy with other work (an idle machine is faster), 400M at full precision (its int8 build is about three times faster). Myles-L on one desktop GPU in 4-bit. It is the escalation tier: an adapter on an open 14B base that reads up to 4,096 tokens and takes the cases the CPU models are unsure of. That base was pre-trained partly on synthetic text from hosted models. Each model's card records its full lineage.

How it compares

Every model answers the same 1,400 decisions, frozen before any comparison: 800 from an open decision dataset (half of them from a shifted distribution), 300 product tasks and 300 everyday tasks. The hosted models answer zero-shot with one fixed prompt.

SystemCorrectCost per million decisionsMedian time
Myles-400M, on one CPU thread91.8%£4500 ms
Hosted LLM, small tier86.9%£900770 ms
Hosted LLM, medium tier89.6%£1,8801,390 ms
Hosted LLM, large tier96.2%£4,1101,870 ms

Myles-400M leads the medium tier by 2.2 points (95% confidence interval 0.4 to 4.0 points) at about one 450th of its cost. Myles is trained for these decisions; the hosted models answer with one fixed prompt. Local cost assumes a cloud CPU at about 3p per hour; hosted prices are list prices converted at £1 = $1.33; hosted times include the network.

Reasoning it was never shown

74%correct on multi-rule problems from domains it never trained on, where chance is 16%
59.7%on an independent public benchmark for decision models, the highest of the open models under 1B parameters we have tested
77.9%on the same benchmark for Myles-L in one forward pass on one GPU

Latest Myles-400M build, September 2026, and Myles-L, October 2026. The reasoning problems are generated with exact answers, and whole domains are held out of training. The benchmark score covers its 231 public tasks, scored with the benchmark's own code.

Confidence you can certify

Myles keeps a decision when its confidence clears a threshold set on held-out data, and hands the rest to a person or a larger model. Each kind of input gets its own threshold. At a 5% target error, the Myles-400M build behind the 91.8% result keeps 73% of decisions, with 2.7% error on unseen test data. At 2%, it keeps 40% with 1.4% error, and at 1%, 16% with 0.5%. Every result lands inside its target.

Organisations can request the Myles-400M  and Myles-L  weights for internal evaluation on Hugging Face, under the Myles Evaluation Licence.

Early access, evaluation licences and pilots.

Contact us

Meet Myles-22M.

A tiny decision model with 22.7M parameters, running in your browser. Give it a message and a question, and it returns a probability for every option in one pass. Nothing you type leaves your device.

Try an example
Ask

Myles-22M is an experimental model and can make mistakes. Please don't rely on it for anything important. Model card