/inference/llama3.2-1b

llama3.2:1b

Measured by 5 nodes on hardware we do not own, and signed by each of them. What machines here actually proved — never what the card claims.

matrix built 10 Sept 2026 · every field verified against the signature of the node that produced it, and re-verifiable by you: each signed figure carries the exact bytes the node signed (signing_payload_b64) and the signature over them, and the ed25519 public key is embedded in node_did. effective_ctx is the context length a node PROVED by recall, never the length a model card advertises
context declared1,31,072what the runtime reported. A node attests it was told this; nobody attests it is true.
context proved by recall8,192the largest size at which a machine could still find tokens planted across the whole prompt — start, middle and end, all of which had to come back.
Advertises 16× the context any machine here could make it use.Not an accusation — it is what happened here, and it is signed.
recalled 8k⌐ declared 128k · unsignedadvertises 16× what it proved
every machine that timed it
18.6 tok/s · unnamed machine13.1 tok/s · unnamed machine9.7 tok/s · x86 · CPU · Hyderabad4.1 tok/s · unnamed machine1.8 tok/s · Apple silicon · Lalganj
1.818.6 tok/s across 5 machines · ×10.5 spreadunnamed machinex86 · CPUApple silicon

This model calls tools correctly and does not stop. A node watched it call the probe tool, then call it again, until the turn budget ran out. Chosen on “supports tools” alone, it will call a tool in production and never come back.

Languages, proved

Each question was put entirely in its own language and script, and the answer had to be right. A model with no coverage cannot parse the question, so answering in English is a failure here, not a pass. 3 proved of 10 asked — a language absent from the probe set was never asked, and is not a claim either way.

Capabilities, measured

Visionprobed, and failed
Shown an image and asked what was in it. The node checked the answer.
Audioprobed, and failed
Given audio and asked to transcribe it.
Reasons about codeprobed, and failed
Asked to predict exactly what a short program prints. This is the half of coding that matters for agent work and that writing valid syntax does not demonstrate: following state through control flow.
Writes codepassed
Asked for a specific function and checked with the Go compiler’s own parser. Nothing is executed — running model-written code to score a model would be a security hole opened for a metric — but syntactic validity is a fact rather than an opinion.
Follows instructionspassed
Given a plain constraint on its output format and checked for exact obedience. This separates an instruct-tuned model from a base one — a base model answers fluently and ignores every instruction, which then looks like a dozen unrelated weaknesses instead of one.
Sustained outputpassed
Asked to produce a long structured answer and checked for length. A model that stops after a couple of hundred tokens cannot write a report, however capable it is otherwise.
Structured outputpassed
The ENGINE was asked to constrain decoding to a schema, and the answer conformed. This is stronger than politely asking for JSON and hoping: the grammar makes invalid output unreachable, which is what an operator actually builds on.
Extended thinkingprobed, and failed
The model emitted reasoning on the engine’s dedicated reasoning channel before answering. Read from that channel rather than by scanning the answer for tags, which is why models that reason are no longer reported as models that do not.
Synthesises sourcespassed
Given three documents with one fact split across them, and required to combine rather than quote. No single document contains the answer, so a model that retrieves without reasoning cannot pass.
Tool callingpassed
The node asked the model to call a function and checked that it called the RIGHT one with a well-formed argument — not merely that it emitted a tool name.
Faithful to tool resultspassed
The model was handed a list through a tool and had to report every item back. Calling a tool and understanding its answer are different abilities: a model can produce a perfect call, receive nine items, and confidently report three. This is the flag that separates them.
Tool loop finishedpassed
Whether the conversation ENDED after the tool call, or the model kept calling until the turn budget ran out. A model that passes tool calling and fails this will call a tool in production and never come back.
Chained tool callsprobed, and failed
The model had to call one tool, read its answer, and call a second with a value derived from it. One call is not agentic work — real specialist jobs are sequences, and a model that never takes the second step looks, from outside, exactly like a specialist silently doing nothing.

The llama family

siblings measured on the same network — compare what each actually proved
signed by this network

What machines here proved

context proved by recall8,192
best throughput49.89 tok/s
machines that measured it5
samples behind those numbers5
regions it was proved in5
Every line traces to a node signature.
from the public record · unsigned

What the runtime says about it

context declared1,31,072
parameters1.2B
parameter count1,23,58,14,432
quantisationQ8_0
familyllama
formatgguf
ollama /api/show; the node attests that the runtime reported these, not that they are true. Attested by did:epn:002408…afbb8a.The only unsigned block on this page. We attest we read these from the runtime — not that they are true. Nobody signs a model card.

Proven across regions

Every machine that measured it

nothing is averaged — the machine is what you are choosing
did:epn:002408…0c5265Bengaluru · 500057 · 0.32.5 · 6 Sept 2026
ctx provedprobe w9lmOdFC1y
18.63tok/s26 samples · 65s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Faithful to tool results Tool loop finished Vision Audio1 languages
did:epn:002408…2e5c02Ranga Reddy · 501202 · 0.32.5 · 24 Aug 2026
ctx provedprobe 47jgHpCTRv
13.09tok/s4 samples · 11s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Faithful to tool results Tool loop finished Vision Audio1 languages
did:epn:002408…9e3c79Bengaluru · 562130 · ollama · 20 Aug 2026x86_64 · intel · 3.0 GiB pooled
ctx provedprobe L/arfOvviN
9.73tok/s94 samples · 629s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Faithful to tool results Tool loop finished Chained tool calls Vision Audio1 languages
did:epn:002408…d68672Bhubaneswar · 751031 · ollama-host · 2 Sept 2026
ctx provedprobe q2W9nrzW5B
4.07tok/s33 samples · 443s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Faithful to tool results Tool loop finished Vision Audio2 languages
did:epn:002408…afbb8aLalganj · 834002 · 0.32.14 · 20 Aug 2026Apple M4 · 10 cores · 16.0 GiB RAM
8,192ctx provedprobe yW7BFtGj9D
1.77tok/s44 samples · 987s
Tool calling Tool loop finished Structured output Extended thinking Vision Audio

Context ladders attempted: did:epn:002408…0c5265 4096 · did:epn:002408…2e5c02 4096 · did:epn:002408…9e3c79 4096 · did:epn:002408…d68672 131072,65536,32768,16384,8192,4096 · did:epn:002408…afbb8a 131072,65536,32768,16384,8192,4096