Myles decision models
One job, done well, on any machine.
Myles reads a situation, a question and the possible answers, and returns a probability for every option in a single pass. It classifies, routes, checks and approves: the decisions inside every AI product.
Myles-400M on our frozen evaluation set. Method below.
The family
| Model | Runs on | Parameters | Typical speed |
|---|---|---|---|
| Myles-22M | Your browser or phone | 23M | 27 ms |
| Myles-150M | A laptop | 149M | 196 ms |
| Myles-400M | The cheapest server you have | 395M | about 500 ms |
| Myles-L | One 24 GB GPU | 14.7B | 97 ms |
| Myles-1B | In training | 1B | — |
Median time for one decision on one CPU thread of a desktop Intel processor: 22M and 150M as 8-bit builds, both timed while the machine was busy with other work (an idle machine is faster), 400M at full precision (its int8 build is about three times faster). Myles-L on one desktop GPU in 4-bit. It is the escalation tier: an adapter on an open 14B base that reads up to 4,096 tokens and takes the cases the CPU models are unsure of. That base was pre-trained partly on synthetic text from hosted models. Each model's card records its full lineage.
How it compares
Every model answers the same 1,400 decisions, frozen before any comparison: 800 from an open decision dataset (half of them from a shifted distribution), 300 product tasks and 300 everyday tasks. The hosted models answer zero-shot with one fixed prompt.
| System | Correct | Cost per million decisions | Median time |
|---|---|---|---|
| Myles-400M, on one CPU thread | 91.8% | £4 | 500 ms |
| Hosted LLM, small tier | 86.9% | £900 | 770 ms |
| Hosted LLM, medium tier | 89.6% | £1,880 | 1,390 ms |
| Hosted LLM, large tier | 96.2% | £4,110 | 1,870 ms |
Myles-400M leads the medium tier by 2.2 points (95% confidence interval 0.4 to 4.0 points) at about one 450th of its cost. Myles is trained for these decisions; the hosted models answer with one fixed prompt. Local cost assumes a cloud CPU at about 3p per hour; hosted prices are list prices converted at £1 = $1.33; hosted times include the network.
Reasoning it was never shown
Latest Myles-400M build, September 2026, and Myles-L, October 2026. The reasoning problems are generated with exact answers, and whole domains are held out of training. The benchmark score covers its 231 public tasks, scored with the benchmark's own code.
Confidence you can certify
Myles keeps a decision when its confidence clears a threshold set on held-out data, and hands the rest to a person or a larger model. Each kind of input gets its own threshold. At a 5% target error, the Myles-400M build behind the 91.8% result keeps 73% of decisions, with 2.7% error on unseen test data. At 2%, it keeps 40% with 1.4% error, and at 1%, 16% with 0.5%. Every result lands inside its target.
Organisations can request the Myles-400M and Myles-L weights for internal evaluation on Hugging Face, under the Myles Evaluation Licence.
Early access, evaluation licences and pilots.