You need a runtime for an agent that will work overnight in CI with nobody watching. You ask in a thread and four names come back, each with a story attached: someone used this at their last job, someone else says the other one is what serious teams run. You open a page of projects sorted by stars and collect a fifth name. Every reply is a name. Not one of them tells you whether the thing picks a session back up after a crash at three in the morning, whether it runs with no terminal attached, whether it keeps the filesystem behind a wall. So you take the name you heard most. You find out in week three.

The back of the box carries the numbers

The front of a food box makes a claim. Natural, wholegrain, the choice of families. The back carries something duller: a table nobody wrote for pleasure, fixed fields in fixed units, in the same order on every box in the aisle. That table is not more honest than the front. It is comparable, which is a different property and a rarer one. It also ends in a barcode, so a phone can read the whole aisle and filter out the one ingredient that would send you to hospital.

A catalogue of harnesses does the same job. Each runtime gets a structured, tagged record carrying the fields that actually decide fit: isolation, headless mode, memory, approval, and what happens after something breaks mid-run. On top sits an interface a machine can query, so a person or an agent asks for candidates matching a set of requirements instead of reading twelve homepages and guessing. None of this is new information. It has been moved into fixed positions where two entries can be held side by side.

A filtered shortlist is not yet a decision

Begin from the requirement, never from the name. An architect asks the catalogue for runtimes with a sandbox, session resume, and support for running under CI, and gets back a shortlist. That is the whole of what it is: entries that survived a filter. Before one of them is wired in, somebody still reads the documentation, the threat model, and the bill.

The distinction matters most where the choice touches what you trust to run alone. Scores and updates can be traceable and still remain signals rather than verdicts. A popularity ranking answers a question about the crowd — how many people had reason to try this — while your question was about an unattended job at three in the morning. Treating the first as proof of fitness is the oldest mistake in the aisle. The sharper version is procedural: an agent can read the catalogue itself, and that is never permission for it to install a dependency because an entry recommended one.

Comparable is only true inside one category

A label helps because everything around it is the same kind of thing. Compare across the aisle and the numbers stop meaning anything. The same holds here, which is why a useful map of this field sorts projects by task, environment, and operating model rather than by vendor or by noise: coding, research, data processing, browser use, orchestration, simulation. A team automating CSV analysis compares data agents with code execution sandboxes and filesystem rules; it does not reach for a general multi-agent framework because that one ships more parts. Word that a project is an agent is not a requirement, and a weekend demo and a production platform do not belong in the same row. What you are choosing, underneath the names, is the loop your agent will live in, and a record goes stale the same way a label does: the recovery behaviour changes, the metadata does not, and the box now lies. Hearsay you can only repeat. A written criterion you can query, check, and correct.