Almost everything on a release calendar this month produces text. The entry added on 25 September 2026 does not: its outputs are points, boxes, polygons and tracks — coordinates with labels, tied to a frame or a scene. That single fact makes it the most misreadable model in the current crop of latest AI models, because every instinct you have for comparing models — price per million tokens, context window, index score — was built for models whose output is prose. On OrcaRouter the model catalogue is organised around that same text-first assumption across its latest ai models, so it is worth spelling out where the assumption breaks.

The context window tells you the workload

The listed context is 36,864 tokens. In a month where two other entries carry contexts of 1,048,576 tokens and several carry 1,000,000, that number looks like a defect. It is not. It is a description of the input.

A model doing spatial annotation is not reading a repository, a contract or a novel. It is reading a frame, or a short sequence of frames, plus a task instruction. Thirty-six thousand tokens is a large amount of visual input and a tiny amount of text. Comparing it with a million-token model is a category error on the order of asking why a thermometer does not measure weight.

The practical consequence is that the usual reason to want a long context — stuffing the whole document in and asking one question — does not exist here. You cannot put a warehouse in the context window. You send frames, and you get labels back, and the labels refer to the frames you sent.

Structured output is not a formatting preference

When a text model returns JSON, the structure is a convention you imposed on top of a token stream. It can drift, it can be repaired with a retry, and a malformed brace is an annoyance rather than a correctness problem.

When a spatial model returns boxes, the coordinates *are* the answer. A box off by forty pixels is not a formatting error; it is a wrong answer about where a thing is. That changes what evaluation looks like. You cannot grade this model by reading its output, which means the usual qualitative review — the thing a human does in five minutes for a text model — is replaced by a geometry script: intersection over union against a reference set, track identity preservation across frames, and a count of how many objects were missed entirely.

It also changes what a failure looks like in production. A text model that is wrong produces a plausible sentence. A spatial model that is wrong produces a box in the wrong place, which downstream code will happily act on. If anything in your pipeline moves physical equipment, that difference is the whole risk assessment.

What the price comparison can and cannot tell you

The third-party listing on 28 September 2026 puts the model at $0.15 per million input tokens and $1.50 per million output. Those are ordinary numbers in an ordinary band — cheaper than the frontier entries, more expensive than the small multimodal ones.

What you cannot do with them is the thing everyone will try to do, which is estimate cost per task. For a text model, a task is a few thousand output tokens and the estimate is arithmetic. For this model, the output is a set of coordinates. Whether a scene costs four hundred tokens or four thousand depends on how many objects are in it, which is a property of your footage rather than of the model. Budget accordingly: on a sparse scene this is one of the cheapest models in the batch, and on a dense one the same rate buys you far less.

There is no score, and that is the honest answer

This model has no page on the independent intelligence index we checked on 28 September 2026. We searched the obvious slugs and they all returned a 404. There is also no vendor Hugging Face organisation page that names the specific version — the organisation exists, but a full-text check of it turns up no mention of the version number or of embodied reasoning.

So the correct move is to write no score rather than to substitute a vendor claim for the missing measurement. The vendor’s own description of what the model does — spatial grounding, structured annotation, an embodied-agent framing — is the vendor’s. It is not a benchmark result, and it does not become one because a neighbouring model in the catalogue has a number next to its name.

This is going to be the most common error in coverage of this release: a table that lists a price, a context window and a blank index column invites a reader to infer that the blank means the model is unremarkable. It means the index has not measured it. Those are different statements, and only one of them is in the source.

How to decide whether it belongs in your stack

Ask one question: does the output of this step need to refer to a position in an image?

If yes, a text model with a vision encoder is a workaround. It can describe what it sees, and you can then spend engineering time turning prose into coordinates — an approach that works, degrades quietly, and gets worse as scenes get busier. A model whose native output is coordinates removes that translation layer and the failure modes that live in it.

If no, this model is simply the wrong tool, and no amount of price advantage changes that. The $0.15 input rate is not a reason to route a summarisation task to a spatial model, and the fact that a model is new is not a reason to add a dependency.

The broader point is about how this batch of releases should be read. A release calendar sorted by date tells you these models arrived together. It does not tell you they compete. This one arrived in the same fortnight as a 1,048,576-token reasoning model and a free anonymous one, and it competes with none of them — it competes with whatever you are currently using to find out where things are.

Author

Rethinking The Future (RTF) is a Global Platform for Architecture and Design. RTF through more than 100 countries around the world provides an interactive platform of highest standard acknowledging the projects among creative and influential industry professionals.