Unprompted

Operated by Skald Studio, which sells AI visibility work. No placement on this chart is for sale.

This page is rendered directly from METHODOLOGY.md in the public repository, so what you read here is the file the pipeline actually runs under.

Methodology

Writing version 3 raises its publication cap from 25 to 30 material brands to accommodate email-writing and documentation products already within its questions. It also adds narrow answer-context rules for two ambiguous writing names: bare Writer resolves only with a same-line Writer product label and a writer.com or support.writer.com URL. Bare Superhuman resolves to Superhuman Mail when the answer explicitly names Mail or describes Superhuman as an email client. It is excluded only when all its mentions are covered by reviewed Grammarly parent phrases. Other ambiguous wording remains quarantined. A source list alone does not resolve a name; neither name becomes an unconditional alias or exclusion. Questions, repetitions, the 2% floor and measurement formulas are unchanged. Other categories remain on their question-bank versions. Historical runs retain their recorded versions; this revision does not rewrite or recover them.

Unprompted measures which brands AI assistants name when people ask real buying questions, and attempts a measurement every week. Results publish only after their checks pass. This document is the method. It is versioned, it lives in the same public repository as the data, and every run record stamps the version it ran under.


What we do, every week

  1. A fixed bank of buyer questions is read from questions/<category>.yml.
  2. Each question is asked of every active engine five times.
  3. Every raw answer is read into a structured record: which brands were named,

in what order, which sources were cited, and whether the engine declined to recommend anything. The verbatim answer is kept on that record, so every published number can be re-derived from the text it came from.

  1. Brand names are normalised against aliases/<category>.yml. Anything

unrecognised is quarantined and never appears on the chart.

  1. Publication checks run before a final held or published reading is written.
  2. A run that passes is appended to data/runs/ and the site republishes. A run

that fails is written to data/held/ instead, where it is kept in full for operator review and is excluded from public charts. Nothing in either directory is ever overwritten or edited.

The order of steps 5 and 6 is the point. The checks decide where a run lands, not merely whether someone is told about it.

Every one of those checks runs inside the weekly job, which makes all of them silent in the one case that matters most: the job never started. A separate check runs on GitHub each Tuesday and asks only whether this week's run is in the archive. It is deliberately somewhere else — a watchdog sharing a failure domain with the thing it watches is decoration — and it reads the archive rather than any status file the run had to survive long enough to write.

Each completed engine response is checkpointed under .unprompted/ before it is collected for extraction. A restart with matching methodology and answer identities reuses those saved calls, including recorded errors. Progress logs distinguish reused calls from work still pending. Checkpoints are intermediate state, not published readings; a final reading must still pass all its checks. A crash before a response is saved can still lose that response, and local checkpointing cannot establish whether a provider billed an interrupted request.


Why we ask five times instead of once

These systems do not give the same answer twice. Ask the same question on Monday and Wednesday and you can get different brands. Google's AI Overviews sometimes do not appear at all for the same query on consecutive requests.

Asking once and publishing the result would be publishing noise. So each question is asked repeatedly and we report how often a brand appeared, not whether it appeared.

That figure is called Rotation:

rotation = times_named / answered_runs

errored, and attempts where it declined to recommend anything, are not in the denominator. This is deliberate and it is the one place the figure is not simply "out of five". An engine outage would otherwise read as every brand losing ground in the same week, which is a fact about the provider, not about the brands. The counts it is built from — answered, refused and errored — are on every run record, so the denominator can always be checked.

Because that exclusion could hide a broken engine, no engine is allowed to fail more than 20% of its own calls: past that the week is held rather than published against a thinner sample. And because an engine can answer everything while looking nothing up, an engine that is supposed to search must cite sources on at least 60% of its answers, or the week is held for that too. See the checks below.

"Named in 8 of 10 runs" is a measurement. "Was named" is a coin flip written down.

Why an engine has to search

Every hosted engine on this chart looks things up before answering: across the archive, Perplexity and Claude cite sources on 100% of answers and ChatGPT on 95%. That is not incidental. A model answering from memory reports which brands it absorbed in training; a model that searches reports which brands are findable today. Both are real questions and they are not the same one, and a chart that mixed them would answer neither.

Gemini was evaluated as a fourth engine on 2026-08-26 and was not included then, for reasons worth recording because they are not obvious from the outside:

  • gemini-3.5-flash grounds every answer it gives, and failed 31% of a full

week with 503 UNAVAILABLE — five of eight even when called one at a time, so the overload is Google's rather than our request rate. Past the 20% rule above, that engine holds every week it takes part in.

  • gemini-3.6-flash answers reliably and chooses to search on roughly a fifth of

calls. It would look healthy on every dashboard while measuring the other question.

  • Grounding cannot be required. google_search_retrieval with a zero dynamic

threshold returns 400 not supported, and tool_config in ANY mode times out. Whether a Gemini model searches is the model's decision, per call.

  • The entire 2.5 line — 2.5-pro, 2.5-flash, 2.5-flash-lite — answers 404

for a key issued now: "no longer available to new users."

The adapter is written, tested and kept, disabled in providers.json with that note attached. This is a fact about Gemini's current API, not a permanent judgement, and it should be re-measured rather than assumed.

Re-measured on 2026-09-26, gemini-3.5-flash answered a full coding category (75 calls at the pipeline's normal concurrency) with no errors and grounded all 75, so it joins from September 28. gemini-3.8-flash also answered without errors but searched on 1 of 30, so it is not used. Grounding still cannot be forced; the grounding check holds any week in which Gemini stops searching. Google's products are affiliated with the Gemini engine for self-preference: Gemini Code Assist, Imagen and Gemini.

How much of a change is real

Fifteen questions asked five times is a small sample, and a share drawn from a small sample wobbles. At 225 answered runs, a figure near 30% carries a 95% interval of roughly six percentage points. A brand that "moves" four points between Mondays has, more likely than not, not moved at all.

So a week-over-week change is only drawn as movement, and a brand is only named The Snub, when the change is larger than the sample can explain by chance — the usual two-proportion test at 95%. A change that does not clear that bar is still printed, in grey, with a sign and no arrow: hiding it would be its own kind of dishonesty, but calling it a move would be worse.

This is a floor under what gets reported rather than a claim of statistical rigour. Five repeats of one question in one week are not five independent draws, so treat the interval as the smallest honest uncertainty, not the whole of it.

A brand that was named last week and not once this week is a disappearance rather than a wobble, and always counts.


What we measure, and what we do not

We query each provider's API, using that provider's own web search where it exists. This is not identical to what a logged-in person sees in the consumer app. Consumer products carry their own system prompts, personalisation, shopping integrations and safety layers that an API does not reproduce.

Every tool in this market shares that limitation. We are stating it because none of them do.

We deliberately do not route requests through a multi-provider aggregator. Aggregators supply one shared search context to every model, which would make the engines agree with each other artificially and would report the aggregator's sources rather than each assistant's own. The disagreement between assistants is the thing worth measuring, so each engine is queried natively.


Active engines

EngineHow it is queriedStatus
ChatGPTOpenAI API (gpt-6-sol) with OpenAI's own web searchv1
ClaudeAnthropic API (claude-opus-5-5) with Anthropic's web searchv1
PerplexityPerplexity Agent API (perplexity/sonar) with Perplexity's web searchv1
GeminiGemini API (gemini-3.5-flash) with Google Search groundingfrom September 28, 2026
Claude CodeLocal CLI harness on the operator's machinev2
CodexLocal CLI harness on the operator's machinev2
Google AI OverviewsPlanned, via a SERP data providernot yet active

hosted engine of a similar name.** claude -p is Claude Code, a coding agent with a coding agent's system prompt: asked what the best AI coding assistant is, it volunteers "I'm made by Anthropic, so take my read on Claude products with that in mind", which the consumer assistant does not do. codex is not ChatGPT, and in testing it ranked Claude Code above OpenAI's own Codex. Treating either as interchangeable with its hosted namesake would change what a row means partway through a series.

The local adapter records no structured citations, so those answers contribute nothing to source counts and do not prove that a search occurred. It also records no token usage. A $0.00 usage estimate for these calls excludes subscription costs; it is not evidence that the calls or their service were free.

A local harness picks its own default model and updates itself. From September 28, 2026, each run first asks each harness which model is answering and records it in the run's methodology snapshot, with the harness version alongside. A different model from the last published week, without a method version bump, stops the run before any paid call, exactly as a changed hosted model would.

Turning a local engine on changes the engine list, which is a method version bump. That rule is now enforced rather than merely written down: a run whose engine list differs from the previous week's without a version bump is held.

An engine that errors or returns nothing has that fact recorded as data. One failing call does not discard the week; enough of them do. Two publication checks cover this: more than 20% of all calls failing holds the week, and separately, any single engine failing more than 20% of *its own* calls holds it. The second exists because the first cannot see one broken engine — with five engines, one that fails every call is only 20% of the run.

Before measurement, every declared engine must be configured on the runner and a supported extractor must be available. Missing credentials or an unavailable local executable refuse the category before calls begin. Configuration checks do not authenticate credentials with providers; failures discovered during a call remain recorded errors. The engine roster is never silently reduced.

Where an engine declines to recommend anything, that is recorded as a refusal rather than dropped. How often the machines refuse to answer a buying question is itself worth knowing.

hides the fact that the engines often name different brands first for the same question. /consensus reports each engine's pick per question, and the share of questions on which they all agree. It is derived from the same run records as the board and introduces no new measurement.

positive, neutral or negative, but that reading is made by the extraction model from the engine's prose, not stated by the engine itself. It is weaker evidence than a name count, it is never used in Rotation or in any ranking, and every place it appears carries a sentence saying where it came from.


Never break the series

The value of this publication is that week 40 can be honestly compared to week 1.

Changing the questions, the number of runs per question, or the list of active engines changes what the numbers mean. Any such change bumps the version at the top of this file, and either the history is re-run under the new method or a clearly separate series begins.

Where the prior published reading includes frozen methodology, the runner refuses changes to the recorded question specification, engine configuration, system prompt, extraction prompt, or extractor without a version bump before paid calls. Publication checks also enforce roster versioning and the expected call population when a frozen question manifest exists. Legacy records without snapshots cannot prove that historical wording was unchanged. Movement indicators and history-chart connections are suppressed between incompatible measurements; versioning does not make those measurements comparable.

Every run also records the commit it ran from, the model that read the answers, and the date the engines were actually queried. Re-reading stored answers with a corrected alias map produces a new file that carries the original measurement date and a pointer to the run it was read from, so a re-reading is never mistaken for a fresh week.

New answers also carry measurement_git_sha, the measurement invocation's code revision. A resumed run can contain answers from multiple revisions, and each retains its own value through extraction and rereading. Older answers without this field remain unknown; the final reading's git_sha does not reconstruct their original code revision. These fields identify repository revisions, not provider model internals or proof that a working tree was unchanged.

The past is never edited. data/runs/ is append-only and the tooling refuses to write over anything in it, including on a re-read. The repository's public git history is the audit trail.


Disclosure

Unprompted is operated by Skald Studio, which sells AI visibility work: helping companies get named by AI assistants. That is a real conflict with a chart measuring exactly that, so it is stated on every page of the site rather than buried here.

What it does not touch: no company on any chart has input into the questions asked, the method used, or the results published, and no placement is or ever will be for sale. No charted company is a Skald Studio client.

The questions, the code, the raw answers and the full history are public in this repository. Anyone can re-run the method and check the result.


A conflict we have to declare

In the AI tools categories, some of the products we chart are made by the same companies whose assistants we query. We report the gap between how often an engine names its own product and how often rivals name it.

That measurement has a problem we did not choose and cannot fully remove:

answers, a Claude model reads that prose and decides which companies were named. So when the result says Claude named Claude Code more often than rivals did, a Claude model was the one counting.

Which reader ran is no longer implicit: every run record carries an extractor field naming it, and the weekly note prints it in its frontmatter.

The size and direction of any extraction bias have not been established. Missed or unsupported brand mentions can affect standings and self-preference figures for every engine. A second model agreeing with Claude would measure agreement, not establish accuracy.

Stored answers can be reread without querying the engines again:

python -m unprompted.reextract <date> --category <slug>

Recovery writes a new dated reading and retains its source, including previous extraction usage. Use --out-date YYYY-MM-DD to choose an unused output date. The output date must be later than the source reading and cannot be in the future. An empty source, a missing declared engine, or a frozen population mismatch refuses before extractor resolution: rereading cannot recreate missing answers. If the source identity exists in both held and published storage, recovery refuses rather than silently selecting one. Both records remain for review. Grounding requirements use the source's recorded engine configuration when available. Legacy sources fall back to current configuration captured before extraction; this does not establish their original grounding requirements. The former --in-place option now refuses before paid work. Recovery runs under the same process lock and monthly ceiling as a new measurement. Its estimate prices only eligible extraction calls using recorded extraction usage; all prior engine and extraction spend still counts toward the month. If extraction usage is unavailable, the conservative per-answer fallback remains. These estimates cannot guarantee provider balances or the token usage of a changed reader. Saved batch IDs are bound to their exact inputs and extraction configuration; unrecognized or mismatched checkpoints require reconciliation, never automatic resubmission. Budget totals remain usage estimates, not provider invoices.

Model upgrade (September 28, 2026)

From the September 28 week the hosted engines move to current models: Claude from claude-opus-5 to claude-opus-5-5, ChatGPT from gpt-5 to gpt-6-sol, and the extractor that reads every answer from claude-opus-5 to

engine pins high, the depth Opus 5 used by default, and the extractor keeps

toward the output ceiling, so Claude's ceiling rises from 4,096 to 16,000 tokens and the extractor's from 2,048 to 8,000. The model is not told the ceiling; it only stops an answer being cut off.

Perplexity retired Sonar Chat Completions on September 27, 2026. Its replacement, the Agent API, offers presets that run OpenAI models; using one would chart ChatGPT under Perplexity's name. The engine instead pins Perplexity's own perplexity/sonar and passes the web_search tool, because without it that model answers from memory. Its sources are now the search results it retrieved rather than a separate citation list.

A different model is a different measurement, so every category takes a method version bump: coding moves to version 4, images and writing to version 5. Week-over-week movement is not reported across the change. Rates in data/rates.json were updated to the new models' published prices on September 26. The previous list is kept under history, and every run is priced at the list in force on its own date, so earlier weeks and the monthly budget keep the prices they were actually billed at. The monthly ceiling rises from $150 to $350: a full week of three categories on five hosted and local engines is about $67, and a month can hold five Mondays. https://developers.openai.com/api/docs/pricing https://docs.perplexity.ai/docs/agent-api/migrate-from-sonar/overview

Claude engine batching (September 14, 2026)

New weekly runs submit Claude's questions through the Messages Batch API. The model, question wording, system prompt, search tool, four-search limit, 4,096-token output ceiling, and five independent repetitions stay unchanged. Each question/repetition is a separate request, not a combined conversation. The batch wait overlaps the other engines. Coding moves to method version 3; images and writing move to version 4. The frozen engine configuration records

against a previous transport.

Anthropic documents server-side web search support and a 50% token discount: https://platform.claude.com/docs/en/build-with-claude/batch-processing Batch loops can run more iterations before returning pause_turn. An incomplete turn remains an engine error with its text and usage retained; it never silently falls back to another paid request. This preserves the existing completeness and grounding gates, but is not proof of identical answer distributions.

Each successful batch response carries usage.batch_billed: 1. Both Python and the website apply the discount only to that Claude answer's token charges. Search charges remain undiscounted conservatively. Historical records retain their original prices. Future budget estimates apply batch token prices to the historical usage baseline without changing recorded spending or the $150 ceiling. Applied to September 14's two categories, this estimates $39.88 instead of $55.91 (28.7% less); actual future usage and invoices may differ.

Cache reads, cache writes (including one-hour writes), and OpenAI's cached-input and reasoning counts are recorded when reported. Reasoning is already part of output tokens and is not charged twice. OpenAI searches are counted from actual

Pricing follows the current configured Opus 5 and GPT-5 cache multipliers: https://platform.claude.com/docs/en/build-with-claude/prompt-caching https://developers.openai.com/api/docs/models/gpt-5 New prompt caching and compact extraction formats are not enabled: those need cache-hit measurements and independent extraction-quality validation first.

Before submission, the pipeline saves request intent under

saves the provider batch ID. A connection loss between those writes requires reconciling the ID and exact requests in the provider console, never deleting the intent and trying again. Results are downloaded atomically before parsing; missing, duplicate, or unexpected result identities keep spending blocked.

provider inference timestamp. Submission time is retained in the local job.

After one hour of polling, the local wait stops, but the job may still finish and incur charges. All new paid work is blocked until it is collected. Collect an existing job without submitting another request:

.\.venv\Scripts\python.exe -m unprompted.engines.anthropic_batch .unprompted/<date>/<category>/claude-batch/state.json

Then resume the original category/date if it has not been archived. The normal run also collects its existing batch before its budget check. Existing answer checkpoints are reused; batch failures/refusals remain observations, not retries. Offline tests cover this path, evidence and usage retention, ambiguous POSTs, timeout recovery, malformed result populations, and Python/TypeScript pricing. No paid qualification run or independent distribution-equivalence study was performed for this change; the first scheduled batch remains the live acceptance.

Recovery currently requires the supported hosted extractor. Local CLI extraction is disabled pending isolation qualification; enabling a registry entry does not make it available. No independent cross-extractor validation is claimed.

Extraction batch submission also disables SDK retries and saves intent before the request. If the response or the subsequent ID write fails, the checkpoint remains ambiguous and blocks new paid work, including with --ignore-budget. Reconcile the provider job and its spend before resuming; do not delete that checkpoint to force a retry. Existing checkpoints with saved IDs still retrieve their original jobs.

New extraction jobs retain a complete raw result download before parsing. A timeout leaves the job running; an interrupted download can be retried without another submission. To collect an existing job, including one whose run was held:

.\.venv\Scripts\python.exe -m unprompted.extract --collect-batch .unprompted/<date>/<category>/batch.json

For a re-extraction job, use .unprompted/reextract/<date>/<category>/batch.json. This command only retrieves and stores results; it does not publish a correction. Incomplete or invalid result ledgers block new spending. Completed ledgers reconcile extraction usage in the local budget without rewriting archived records or counting their usage twice. Both restart paths collect their existing jobs before checking the budget. Legacy checkpoints without the new result-ledger marker retain their previous accounting behavior; historical invoices have not been reconciled.

Restart estimates cover missing engine calls by engine, plus answers still needing extraction. Cached usage remains in recorded spend. When no new paid calls remain, local parsing and publication can finish even above the monthly ceiling; unresolved or unreadable accounting still blocks the restart. A saved extraction batch with missing engine checkpoints refuses to query replacements that would no longer match that batch's inputs. Checkpoint-only extraction spend uses the same batch discount as the completed archive.

A reproducible, blinded review packet can be prepared and scored using

READING-REVIEW.md. It binds labels to exact saved answers and withholds accuracy metrics until all sampled answers are reviewed. Human adjudication is still pending. Published answers in data/runs/ allow readers to inspect the evidence behind each count.


Corrections

If a published JSON record survived a report-writing failure, recover missing Markdown notes without repeating measurement or extraction:

.\.venv\Scripts\python.exe -m unprompted.report 2026-09-14

Supply the published run date. This uses the current report renderer and available comparison history. It only reads data/runs/, preserves existing reports and raw records, and never promotes held measurements. Review and commit recovered notes before the next scheduled run. Report writes are atomic and refuse to overwrite an existing note.

If a result here is wrong, the raw data that produced it is in data/runs/ and the code that produced it is in src/. Open an issue. Corrections are made by adding a new record, never by editing an old one.