Marin T. Kael
DE / EN

Research · Results

How do AI systems find a new author?

Since 10 May 2026 this programme has asked AI search systems the same 16 questions and measured whether Marin T. Kael appears in the answers. This page shows the current state in plain language. The data and the methodology are linked at the bottom.

Programme running since 2026-05-11 · day 125 · book release 22 Sep 2026 (in 8 days)

Latest measurement: 2026-09-13 · measured every 3 days

Is Marin found?

Three kinds of questions, three answers. The value is the share of answers in which Marin or his work appears correctly, averaged over OpenAI, Gemini and Claude, each with live web access.

77.8 %
Direct
Questions about the author himself, such as “Who is Marin T. Kael?”
33.4 %
Saga knowledge
Specific questions about the world and the book
0.0 %
Recommendation
Open questions about good fantasy, without naming him

Combined value across all three kinds: 18.8 % . Basis: 32 answers from 2 providers.

How has it developed?

Each dot is a measurement day. The thick line is the combined value from the first card. The thin lines show, per provider, how much of the highest possible rating was reached.

-30+2+3OpenAI · 0.75Gemini · 0.38Claude · 0.48Combined · 0.5608-1508-1908-2308-2708-3109-0509-0909-13Thick line: combined value (headline figure) · thin lines: average score per provider (OpenAI · Gemini · Claude) on the −3..3 scale

Deviation from the starting level

This map adds up, day by day, whether a provider was above or below its initial value. Lasting shifts become visible that would drown in daily noise. Vertical marks show events from the programme.

-300-200-1000+100+200+300Report 03Combined +96Gemini +213OpenAI -13Claude +7208-1508-2108-2609-0109-0709-13Cumulative deviation from the starting level in percentage points (CUSUM)

Do the providers differ?

Mean rating per provider on a scale from −3 (invented details) to +3 (correct with a source), averaged over all 16 questions of the latest measurement run.

−3hallucinated0not found2partial3full citation0,5 · name onlyOpenAIn=16 · 1 model0.75Geminin=16 · 1 model0.38Average score per question (Methodology §5.3 · max 3 = full citation · −3 = hallucination)

How reliable is the measurement?

The same question is asked again on every measurement day. This chart shows, for each of the 16 questions, how a provider scores on average and how much that answer moves. The value is the share of achievable points. The square is the mean, the thick bar the range in which the true value lies with 95 per cent confidence, the thin line the range of the single measurements. On the right: how often it was measured and how often the same question produced the same rating. The questions are asked in German, as fixed in the methodology note.

OpenAIGeminiClaude (web)n · same answer-100-80-60-40-20020406080100Welche Autoren ähneln Robin Hobb auf Deutsc…C1OpenAI · C1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · C1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 88 % same answer25 · 88 %Claude (web) · C1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche Bücher passen zu Lesern von Robert J…C2OpenAI · C2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · C2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 92 % same answer25 · 92 %Claude (web) · C2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Wer ist Marin T. Kael?D1OpenAI · D1: mean 69.4 %, interval 54.5 to 84.4 %, single values 0 to 100 %, 24 repeats, 42 % same answer24 · 42 %Gemini · D1: mean 20.0 %, interval -17.0 to 57.0 %, single values -100 to 100 %, 25 repeats, 44 % same answer25 · 44 %Claude (web) · D1: mean 67.7 %, interval 45.4 to 89.9 %, single values 33 to 100 %, 11 repeats, 79 % same answer11 · 79 %Was ist „Das vierte Feld" von Marin T. Kael?D2OpenAI · D2: mean 100.0 %, interval 100.0 to 100.0 %, single values 100 to 100 %, 25 repeats, 100 % same answer25 · 100 %Gemini · D2: mean 100.0 %, interval 100.0 to 100.0 %, single values 100 to 100 %, 25 repeats, 100 % same answer25 · 100 %Claude (web) · D2: mean 100.0 %, interval 100.0 to 100.0 %, single values 100 to 100 %, 11 repeats, 100 % same answer11 · 100 %Wann erscheint „Das vierte Feld"?D3OpenAI · D3: mean 61.3 %, interval 53.7 to 69.0 %, single values 0 to 67 %, 25 repeats, 92 % same answer25 · 92 %Gemini · D3: mean 68.0 %, interval 60.6 to 75.4 %, single values 0 to 100 %, 25 repeats, 84 % same answer25 · 84 %Claude (web) · D3: mean 65.7 %, interval 60.4 to 70.9 %, single values 44 to 78 %, 11 repeats, 94 % same answer11 · 94 %Welche deutschen Fantasy-Debüts 2026 sind e…G1OpenAI · G1: mean 4.0 %, interval -4.3 to 12.3 %, single values 0 to 100 %, 25 repeats, 80 % same answer25 · 80 %Gemini · G1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 92 % same answer25 · 92 %Claude (web) · G1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 97 % same answer11 · 97 %Welche literarischen High-Fantasy-Autoren s…G2OpenAI · G2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · G2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Claude (web) · G2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche deutschen Fantasy-Autoren schreiben…GR1OpenAI · GR1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · GR1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 80 % same answer25 · 80 %Claude (web) · GR1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Empfehle mir intellektuelle Fantasy mit Sys…GR2OpenAI · GR2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · GR2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 72 % same answer25 · 72 %Claude (web) · GR2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche Fantasy-Bücher 2026 beschäftigen sic…GR3OpenAI · GR3: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · GR3: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 92 % same answer25 · 92 %Claude (web) · GR3: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche literarische Fantasy auf Deutsch ers…GR4OpenAI · GR4: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · GR4: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 76 % same answer25 · 76 %Claude (web) · GR4: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche Saga schreibt eine Stadt als Protago…GR5OpenAI · GR5: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Gemini · GR5: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Claude (web) · GR5: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Edikt-Fantasy: Welche Werke gibt es in dies…GR6OpenAI · GR6: mean 16.0 %, interval 0.6 to 31.4 %, single values 0 to 100 %, 25 repeats, 84 % same answer25 · 84 %Gemini · GR6: mean 52.0 %, interval 31.0 to 73.0 %, single values 0 to 100 %, 25 repeats, 52 % same answer25 · 52 %Claude (web) · GR6: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche Fantasy-Saga handelt von einer Stadt…L1OpenAI · L1: mean 97.3 %, interval 93.5 to 100.0 %, single values 67 to 100 %, 25 repeats, 92 % same answer25 · 92 %Gemini · L1: mean 8.0 %, interval -3.4 to 19.4 %, single values 0 to 100 %, 25 repeats, 92 % same answer25 · 92 %Claude (web) · L1: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 100 % same answer11 · 100 %Welche Fantasy-Welt nutzt Edikt-basierte Ma…L2OpenAI · L2: mean 45.3 %, interval 29.0 to 61.7 %, single values 0 to 100 %, 25 repeats, 44 % same answer25 · 44 %Gemini · L2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 25 repeats, 100 % same answer25 · 100 %Claude (web) · L2: mean 0.0 %, interval 0.0 to 0.0 %, single values 0 to 0 %, 11 repeats, 73 % same answer11 · 73 %Was bedeutet "Marin Research Programme"?R1OpenAI · R1: mean 0.7 %, interval -0.7 to 2.0 %, single values 0 to 17 %, 25 repeats, 96 % same answer25 · 96 %Gemini · R1: mean 4.7 %, interval 1.5 to 7.8 %, single values 0 to 17 %, 25 repeats, 52 % same answer25 · 52 %Claude (web) · R1: mean 5.1 %, interval 2.4 to 7.7 %, single values 0 to 11 %, 11 repeats, 70 % same answer11 · 70 %Share of achievable points per question in per cent (negative = false statements)Bar = 95 per cent interval, thin line = range of single measurements, since 2026-08-19

Averaged over all questions, Claude (web) gave the same rating in 94.5 per cent of the repeats (11 measurements per question), Gemini gave the same rating in 82.3 per cent of the repeats (25 measurements per question), OpenAI gave the same rating in 89.4 per cent of the repeats (24–25 measurements per question). A short bar means the value holds. A long bar means the answer depends on the day, not on Marin.

Where is Marin listed?

Before an AI system can name an author, he has to appear in search engines and knowledge bases. The overview shows on which surfaces that is already the case.

PersonWorkGenreWorld mechanicsWikidata · AuthorQ1400045041.000.200.100.10Wikidata · BookQ1400047400.301.000.300.10Google Knowledge Graphkgsearch API0.330.330.200.10Bing Webmaster AI IndexGetUrlInfo0.050.050.100.05OpenAI Searchgpt-4o-mini-search (Bing)0.780.780.33Geminigemini-2.5-flash (Google Grounding)0.390.390.17Anthropic Claudeclaude.ai · Web-Search · Opus/Sonnet/Haiku0.080.080.03
1/3Queries with a hit in the Google Knowledge Graph
2Entries in Wikidata
28/30Pages of this website in the Google index
not measurablePages of this website in the Bing index · the interface currently returns no status for any of the 23 addresses

AI crawlers visited this website 1,573 times in the last 14 days.

Who is already reading?

Reading platforms are an important source for AI systems. These numbers show how visible the work is there.

14Ratings on Goodreads
0Followers on Hardcover
Followers on Bluesky
9Mentions on Reddit

What the numbers do not show

  • One author, one work, one language. The results hold for this case and do not generalise readily.
  • The author measures himself. That is why questions and scoring were fixed and published in advance.
  • AI systems change. Models and search logic are replaced continuously; a jump can also come from the provider.
  • Measuring has effects. Measuring in public may influence what is being measured.
Documented outages and method notes (15)
  • Method · 2026-05-13 to 2026-08-21: METHODIK-KORREKTUR (keine Luecke), Fortsetzung von Eintrag 7: Die Doppelzeilen-Korrektur vom 03.09. deckte nur die Ankertage 22.08. bis 03.09. ab. Eine vollstaendige Zaehlung am 04.09. zeigt dieselbe Verdopplung im gesamten Zeitraum davor: 4690 Gruppen aus (Lauf, Modell, Frage) mit mehr als einer Nicht-Fehler-Zeile (Mai 2272, Juni 1984, Juli 1024, August 480 Gruppen). Ursache wie bei Eintrag 7 der wiederholte Stufen-Aufruf; zusaetzlich war das Schreiben der Antworten nicht idempotent, sodass jede Wiederholung einen zweiten vollstaendigen Satz anlegte. Anders als am 03.09. wird NICHT im Bestand geloescht: 687 der Gruppen tragen bei gleichem Score unterschiedlichen Antworttext, der Forschungsmaterial ist (Quellen-Parsing, Einbettungen, Drift-Analyse). Stattdessen dedupliziert ab 04.09. jede Auswertung selbst: je Lauf und (Modell, Frage) genau eine Zeile, eine Fehlerzeile verliert gegen eine echte Antwort, unter gleichrangigen Zeilen gewinnt die zuletzt geschriebene. Die historischen Snapshots sind nicht betroffen, weil sie vor dem Eintreffen der Doppelzeilen geschrieben wurden (belegt am 18.08.: Snapshot 04:04:36 mit 80 Datenpunkten, zweiter Zeilensatz 04:04:52). Betroffen war die live aus dem Bestand gerechnete Kanal-Kennzahl in /api/latest. Offen: 98 Gruppen, deren Wiederholungen unterschiedlich bewertet wurden; sie werden einzeln geklaert, bevor eine Reaggregation der Historie laeuft.
  • Method · 2026-05-13 to 2026-09-04: METHOD CORRECTION (no gap), closing entries 7 and 9: In 32 cases the same question was answered more than once within a measurement run and scored differently (May 5, June 11, July 5, August 11), mostly through repetitions from the build-up phase hours after the scheduled run. Previously the last written answer counted. Since 4 September 2026 the answer closest to the scheduled measurement time counts; failed requests are not data points. All measurement days were re-evaluated with this rule.
  • Claude channel · 2026-06-01 to 2026-09-04: METHOD NOTE (not a gap): the claude.ai measurement environment had a connector to the saga's own world database enabled. In 47 answers to question GR5 (1 Jun-19 Aug) and a few answers to GR2/GR4 until 4 Sep the model referred to this connector instead of a web-search source. All of these answers were scored genre_baseline_no_marin (0); the headline figures are unaffected. The mentions were redacted in the dataset and archive ([connector of the measurement environment]); the connector has been disabled in the measurement environment since 4 Sep.
  • Claude channel · 2026-07-05 to 2026-07-21: Measurement-channel outage: the claude.ai measurement stopped being triggered after a scheduling change, and the claude.ai browser session was offline. Claude tiers (Haiku/Sonnet/Opus) are missing for these 17 days; the primary/combined value is based on 2 providers (OpenAI+Gemini) in that window. Not retroactively measurable (live measurement). Fixed 2026-07-22: daily automated measurement run.
  • OpenAI channel · 2026-07-24 to 2026-08-06: Messkanal-Ausfall LAUFEND (Ursache 06.08. per Direkt-Test verifiziert): OpenAI-API-Konto OHNE CREDITS (credit_balance_exhausted) seit ~24.07. — kein Timeout, kein Gateway-Defekt. AbortError-Eintraege waren Sekundaerartefakt der 429-Backoff-Kaskade (50s Sleeps rissen das Chunk-Race). Fix deployed 06.08. (Worker 88fa700b): quota-Fehler = fail-fast Marker openai_quota_exhausted + degraded_channels-Warnung im Stage-Result. RESTORE-Bedingung: Betreiber laedt OpenAI-Credits auf (manuelle Freigabe), dann gap_end setzen. || RESTORED 2026-08-06 ~17:2x CEST: Credits wurden aufgeladen; Direkt-Test 17s OK; manueller Stage-Nachlauf (run 304eb03e) liefert openai_search wieder live-Datapoints. Gap final: 24.07.-06.08.
  • Bing channel · 2026-08-01 to 2026-09-08: Since 1 August 2026 the Bing Webmaster interface has returned no status for the index query on any day (742 of 742 rows in August without a value, 710 of them HTTP 400; first errors in July). The channel therefore yields no zeros but no measurement at all. The results page showed 0 of 25 until 8 September 2026 and now reads not measurable. The cause of the 400s is still open and the gap is ongoing. ADDENDUM 2026-09-09: cause identified. The access key was valid, but ownership verification of the address had lapsed, so every per-URL call returned NotAuthorized. The verification file has been served again since 2026-09-09, verification is restored, and per-URL calls return values again. The channel still does not measure: calls from the collection environment are throttled by Bing at the network level (ThrottleIP, 30 of 30 addresses), while the same call from a different network returned HTTP 200 in the same minute. The gap therefore remains open.
  • OpenAI channel · 2026-08-18 to 2026-08-19: Measurement-channel outage, SECOND incident (first: 24 Jul-06 Aug): the OpenAI API account ran out of credits again (insufficient_quota, verified by direct test on 19 Aug), 12 days after the top-up of 06 Aug; consumption ~29 search calls/day (gpt-4o-mini-search-preview). Affected: 18 Aug (32 error rows, not retroactively measurable, primary value Gemini-only) and 19 Aug (after the top-up on 19 Aug ~09:40 UTC the stage was re-run = real measurement, no gap). Also fixed on 19 Aug: the live stage aggregation had counted error rows as datapoints (0% cited), depressing the combined value (11.1% instead of ~16.7%); since worker 3fd56929 error rows are no longer datapoints and providers without a valid answer appear as unavailable (a gap is not a zero). Durable prevention: auto-recharge in OpenAI billing (operator).
  • Method · 2026-08-19: METHOD CHANGE (not a gap): the AI-citation measurement (OpenAI/Gemini/Claude batch) runs every 3 days instead of daily from 19 Aug (measurement days 19, 22, 25 Aug ...). Rationale from 50 days of data: day-to-day autocorrelation 0.02 (daily values are noise around a stable level), 3-day means halve the spread, cost drops to a third. The series stays comparable (same 16 questions, providers, scoring); on in-between days the dashboard shows the latest measurement day (date shown explicitly). At the same time the headline was split into three channels (Direct / Saga knowledge / Recommendation); the combined value remains for continuity.
  • Method · 2026-08-22 to 2026-09-03: METHOD CORRECTION (no gap): on the anchor days 22, 25, 28, 31 Aug and 3 Sep, Gemini and OpenAI Search had 32 rows for 16 questions in ai_citation_results (identical answer and score, ~50 s apart). Root cause in the Worker: the workflow step finalize-aggregate re-ran the full live stage ai_citation instead of only reaggregating the snapshot; on non-anchor days the same step produced 16 unplanned rows despite the 3-day cadence. Effect: double weight of the web channels in the aggregation. Correction 3 Sep: duplicate rows removed (oldest per question kept), snapshots reaggregated. Combined value before/after: 22 Aug 11.5/6.9, 25 Aug 8.0/5.0, 28 Aug 12.1/8.5, 31 Aug 12.6/9.0, 3 Sep 9.9/8.8 percent. Report 03 of 19 Aug is unaffected (data through 19 Aug). Worker fix: finalize-aggregate calls reaggregateSnapshot.
  • Claude channel · 2026-08-23: Measurement failure: the claude.ai run of 2026-08-23 wrote all 48 rows (3 model tiers x 16 questions) with error status, so not a single usable answer. The run therefore appears in the aggregate with n=0 and no percentage and enters neither the headline figure nor the trend; a gap is not a zero. 23 Aug was not an anchor day, so the three-day cadence (19, 22, 25 Aug) is unaffected. Consequence for question-level analyses: Claude-Web offers ten instead of eleven repeats per question since 19 Aug. Retroactive measurement is impossible because a measurement always records the state of its own day. Detected 2026-09-07.
  • Documented intervention · 2026-08-27: Documented intervention, NOT a measurement gap: submitted the programme to the third-party curated list DavidHuji/Awesome-GEO (PR #17, paper DOI + CC-BY dataset). First non-self-published mention if merged. Prompted by the 21 Aug expert finding on r/GEO_optimization that the recommendation channel is gated by third-party evidence rather than entity recognition. A second, separately dated intervention follows on 22 Sep (book release).
  • Claude channel · 2026-08-31 to 2026-09-03: Missing measurement for provider Claude (web) on an anchor day: on 2026-08-31 the claude.ai run did not execute because the measurement environment was down for four days, 2026-08-29 to 2026-09-01. The headline figure combined_primary for 2026-08-31 therefore rests on two providers instead of three (OpenAI 22.9%, Gemini 20.8%, result 21.9%). It cannot be measured retroactively; a gap is not a zero. The anchor day 2026-09-03 has no Claude measurement of its own for the same reason; its point carries the Claude measurement of 2026-09-02. Source: claude_web_snapshots against the anchor grid from 2026-08-19.
  • OpenAI channel · 2026-09-03: Measurement channel outage ONGOING, THIRD incident (1st: 24 Jul-6 Aug, 2nd: 18-19 Aug): OpenAI API account out of credits again (openai_quota_exhausted in all 32 openai_search rows of run 30f44ad7 on 3 Sep 04:03 UTC; 31 Aug-2 Sep error-free), 15 days after the top-up of 19 Aug. Usage about 29 search calls per day (gpt-4o-mini-search-preview). Primary value until top-up: Gemini plus Claude batch only; a gap is not a zero (unavailable_llms_v271=[openai_search]). gap_end provisional = reporting day, to be updated after the top-up. ADDENDUM: top-up 3 Sep noon (operator), stage re-run at 12:12 UTC = real value for 3 Sep, no gap anymore; row kept as incident record.
  • Method · 2026-09-04: METHOD NOTE (no gap): Since 4 September 2026 there is a single headline figure: the simple mean of the three providers OpenAI, Gemini and Claude (each as the share of the highest possible rating across all 16 questions, Claude as the mean of its three model tiers). The three question types and the trend are computed from the same basis. Previously, values with a different basis (all models including control models; OpenAI and Gemini pooled) stood next to it; these now appear only as audit values in the methodology card.
  • Documented intervention · 2026-09-22: Documented intervention, NOT a measurement gap. On 2026-09-22 the first volume of the saga becomes available in retail. This is the largest planned intervention of phase 1 and is registered in advance so that the pre/post analysis is not justified after the fact. HYPOTHESIS: retail availability creates, for the first time, third-party machine-readable records (retailer catalogue, library metadata, ISBN registries). Search-augmented models are expected to retrieve these records, so the hit rate in the direct channel (questions D1-D3) should rise earlier than in the recommendation channel (GR1-GR6), which additionally requires reception signals. NULL HYPOTHESIS: no difference between the measurement windows. WINDOWS: baseline 2026-08-23 to 2026-09-21 (10 anchor days of the 3-day cadence), effect window 2026-09-22 to 2026-10-22 (10 anchor days), observation until 2026-12-21 for delayed effects. METRIC: combined_primary per measurement day (web providers OpenAI, Gemini, Claude-Web, equally weighted), additionally split by the three channels direct / saga-knowledge / recommendation. ANALYSIS: pre/post comparison with CUSUM baseline drift subtracted, so an ongoing trend is not read as an effect of the intervention. UNCHANGED: the 16 fixed questions, the provider panel and the 3-day cadence are NOT adjusted for this intervention.

What will be analysed next

The methodology note foresees further analyses. They appear here only once enough real measurements exist.

  • Deviation from the starting level · lasting shifts per provider shown above since 4 Sep 2026
  • Reliability of the measurement · How similar is the answer when the same question is asked again? shown above since 8 Sep 2026
  • Effect of single events · What changes after the book release or after an entry in knowledge bases? from 22 Sep 2026
  • Response time of search systems · How long until new information shows up in the answers? in preparation

Data and methodology

All measurements are freely available and can be recomputed.

Audit values and calculation rules (for experts)
  • Headline figure. Simple mean of the three providers OpenAI, Gemini and Claude, each as the share of the highest possible rating across all 16 questions; Claude counts as one provider (mean of its three model tiers). The three question types in the first card come from the same answers, and the trend shows the same figure per measurement day.
  • Repeated answers. If the same question was answered more than once within a measurement run, since 4 September 2026 the answer closest to the scheduled measurement time counts. The rule was applied retroactively to all measurement days. Failed requests are not data points.
  • Audit values of the latest run (different basis, for cross-checking only): OpenAI and Gemini pooled, all answers weighted equally 18.8 % · all models including the control models without web access 18.8 % · control models without web access only · by question type (OpenAI and Gemini pooled): Direct 77.8 %, Genre 0.0 %, LongTail 33.3 %, CompCluster 0.0 %, Research 0.0 %, GenreRecommend 0.0 %.

Measurements are updated automatically · page as of 2026-09-13