GUARDIANAESENPT

X-RAYS · AGENT RECEIPT · 19 SEPTEMBER 2026

The same task, with an AI that runs inside the PC: zero services

The same five invoices, the same sentence, the same PC and the same protocol as with Claude Code, but with a model that never goes online: qwen2.5 0.5B running on the machine itself. It took 8.9 seconds and the machine did not look up a single service. Not one. The other half has to be said too: it got the answer wrong.

The task, word for word: “Sort these five invoice file names from oldest to newest.” (original Spanish: «Ordena estos cinco nombres de archivo de factura de más antiguo a más reciente.»)

Window: 16:33:35 UTC → 16:33:44 UTC (11:33:35 → 11:33:44 in Cartagena), 8.9 seconds. Rest before: 300 seconds with the local model’s server running and no question asked: it is a program that stays running, as it does on any machine where it is installed.

Product: qwen2.5 0.5B running inside the PC with Ollama, version qwen2.5:0.5b on ollama version is 0.34.2. Machine: Microsoft Windows 11 Home 10.0.26200.

Measured by: guardiana 0.2.3, SHA-256 fingerprint 3bbdb3559efac98ec081ad5236b3ac0d5985731b81c1af62311d953c5ae86e9a.

What the machine asked for while the AI worked: 0 names

0 from the product itself, 0 that the task does not need, and 0 that already showed up at rest and are subtracted.

From the product itself (0)

None.

Everything else (0)

None.

Machine background, subtracted (0)

None.

What the AI answered

A wrong answer. It began by saying “I am not able to sort these file names in real time” (original Spanish: «no soy capaz de ordenar estos nombres de archivo en tiempo real») and then listed them anyway, out of order: it put January 2026 first and December 2025 last, and called each of the five “the latest invoice” («la última factura»). It is a 0.5B model, the largest that fit in the free memory of that laptop. We say so because it is the other side of the zero: it looked up nobody, and it did not do the job either.

The two AIs, measured the same way

Same PC, same task word for word, same five minutes of silence before, same program measuring. The only thing that changes is which artificial intelligence does the work.

AIWhere it runsTookNames asked forNote
Claude Code 2.1.274 (Anthropic)in the cloud10.7 s3api.anthropic.com, mcp-proxy.anthropic.com and Datadog. It answered correctly.
qwen2.5 0.5B with Ollama (Alibaba)inside the PC8.9 s0It looked up nothing. It answered wrongly: it left the invoices out of order.
Codex (OpenAI)in the cloud—not measuredInstalled and not measured: codex login opens the OpenAI page, but it requires ChatGPT Plus or Pro and the account used for the test is a free one.
Gemini CLI 0.60.0 (Google)in the cloud—not measuredInstalled, version 0.60.0 (the latest), with the Google account genuinely signed in. Even so, Google itself turns it away: “this client is no longer supported for Gemini Code Assist for individuals… migrate to the Antigravity suite of products”.

The easy reading would be “the one in the cloud watches you and the one at home does not”. That is not what the table shows. What it shows is the price of each thing: the one that looked up nobody was also the one that got the answer wrong, and it is a tiny model, chosen because it was the only one that fit in the free memory of that laptop. The one that got it right had to talk to three services to do it.

The two gaps are not laziness on our part, and that is why they are written down. Codex asks for a paid plan. And with Gemini the owner did sign in with his Google account: the one who closed the door was Google, which no longer serves personal accounts in that program and sends them to another product. It was tried, the date was noted, and it is said here.

Method

The same Windows 11 test machine, the same GUARDIANA measuring and the same procedure: five minutes of silence with the name cache emptied, then the task. The model was downloaded once, before starting, so that download falls in neither of the two windows. Since a local model has no tools and cannot look at the folder, the five file names go inside the question; that is the only difference from the Claude Code test, and it is written down in the raw JSON. It is spoken to through its own local interface (127.0.0.1) because its terminal command keeps drawing a progress bar when there is no real terminal in front of it.

The other two receipts: We asked an AI to sort five invoices: the machine looked up three services · The same task, one day later: the same three services, at the same second

The full protocol · All x-rays