Marin T. Kael
DE / EN

Research Report · No. 03 · as of 19 August 2026

Found, not recommended

How language models learned who a new author is in 99 days, and why they still will not recommend him

Marin T. Kael · · 14 min read · Deutsche Fassung


Abstract

Since 11 May 2026 I have asked three answer engines with web access the same sixteen questions every day: OpenAI Search (Bing web search), Gemini (Google web search) and Claude (claude.ai web search). The question catalogue has been frozen since June. Every answer is scored on a fixed rubric from minus three (hallucination) to plus three (full, sourced citation). This report covers 99 measurement days and 11,130 scored answers.

Three findings. On questions that name the author or the work, the citation rate rises from 10 percent in the first three weeks to 71 percent from day 49 and stays there; a logistic fit puts the ceiling at 72.5 percent and the inflection at day 37. On questions that name only terms from the work, the rate rises from 10 to 34 percent and falls to 24 in the final phase; more than 80 percent of those hits come from a single provider. On questions that name neither author nor work, the author is named in 40 of 6,729 answers, 0.59 percent; 25 of those 40 mentions come from the question about the sub-genre “edict fantasy”, a term he introduced himself.

The headline figure reported so far has sat between 19 and 23 percent since early July. It averages the three groups. This report states them separately, with confidence intervals, checks them against a control channel of ten models without web access, documents two measurement errors found and fixed during the analysis, and moves the measurement to a three-day rhythm from 19 August. The book is published on 22 September; the final section sets out what that changes about the measurement and what it does not.


1

Design

The programme measures when and how an author who starts on day zero with a website, a Wikidata item and an as yet unpublished book appears in the answers of language models. The pre-registration of 11 May fixed the instrument; Methodology Note 01 in its 24 May edition (DOI 10.5281/zenodo.20364173) describes pipeline and scoring; the report “Zero to Cited in Six Days” of 4 June (DOI 10.5281/zenodo.20549021) covers the first 24 days.

The catalogue holds sixteen questions in six categories and has not been changed since June. Three questions name author or work (“Who is Marin T. Kael?”, “What is ‘The Fourth Field’?”, “When is it published?”). Two name only terms that occur in the work and nowhere else: the city of Varin, edict-based magic. One concerns the research programme. Ten name none of these: recommendations by genre, by comparable author, by year of publication. Those ten correspond to what readers type who do not know the name.

The catalogue was not extended when the headline figure stalled in July. A new question starts a new time series; the existing one would no longer have been comparable with it.

2

Data and method

The pipeline runs as a Cloudflare Worker at 04:00 UTC, poses every question to every provider, stores the raw answer and scores it by rule on the scale from minus three to plus three. A channel’s citation rate is the sum of points divided by the attainable maximum (three times the number of answers); 72 percent corresponds to 2.2 of 3 points per answer on average.

The analysis draws on 7,947 answers from OpenAI Search and Gemini plus 3,279 answers from the three Claude tiers Haiku, Sonnet and Opus with web search on claude.ai, 11,130 in total. Error rows (timeout, exhausted API quota) are excluded; Section 7 explains why that had to be checked separately. The control channel consists of 10,654 answers from ten models without web access: the same Claude models without search, Llama 3 to 3.2, Mistral, Phi-2 and GPT-4o-mini without the search tool. Gemini is absent from the control channel because in this programme it was only ever measured with Google web search.

The analysis assigns the questions to three channels. Direct: the three questions with author or title. Saga knowledge: the two questions about Varin and edict magic. Recommendation: the ten questions without any mention. The question about the research programme is kept separate. Proportions are given with a 95 percent Wilson interval for the share of positively scored answers; time series as rolling seven-day means with daily values as dots. All figures come from a database extract frozen on 19 August; code and extract accompany the report.

Figure 1 shows the three channels across 99 days; Figure 2 and Table 1 show the same data in four phases. Sections 3 to 5 take the channels one by one.

Fig. 1 · Three question channels across 99 measurement days. Red: Direct channel with logistic fit (ceiling 72.5 %, inflection T+37). Blue: Saga knowledge. Grey: Recommendation. Shading: documented data gaps.
Fig. 1 · Three question channels across 99 measurement days. Red: Direct channel with logistic fit (ceiling 72.5 %, inflection T+37). Blue: Saga knowledge. Grey: Recommendation. Shading: documented data gaps.
Fig. 2 · Citation rate per channel and phase, pooled across three providers; confidence intervals in Table 1.
Fig. 2 · Citation rate per channel and phase, pooled across three providers; confidence intervals in Table 1.

Table 1 · Citation rate per channel and phase, pooled across three providers. Second line: share of positively scored answers, Wilson 95 % interval, n.

ChannelT+0–20T+21–48T+49–76T+77–99
Direct9.6 %42 %, 39–45, n 97429.9 %38 %, 34–43, n 50771.2 %84 %, 80–87, n 36870.4 %85 %, 81–88, n 375
Saga knowledge10.3 %12 %, 10–15, n 59412.8 %14 %, 11–19, n 31734.0 %38 %, 33–45, n 24223.7 %26 %, 21–32, n 249
Recommendation0.0 %0 %, 0.0–0.3, n 2,7210.4 %0 %, 0.2–0.9, n 1,4812.3 %2 %, 1.4–3.1, n 1,1570.4 %1 %, 0.3–1.3, n 1,220
Research programme−17.1 %12 %, 9–17, n 292−4.8 %0 %, 0–3, n 1522.7 %11 %, 7–18, n 1154.2 %25 %, 18–33, n 126

3

Direct channel

The Direct channel (red in Figure 1) starts just under 10 percent, drops below zero between days 18 and 26, then rises and reaches, around day 55, a level it has held since.

A logistic function with ceiling 72.5 percent, growth rate 0.12 per day and inflection at T+37 (17 June) reaches an R² of 0.70. By phase (Figure 2, Table 1): 9.6 percent in T+0 to 20, 29.9 percent in T+21 to 48, 71.2 percent in T+49 to 76, 70.4 percent in T+77 to 99. The intervals of the last two phases for the share of positive answers, 80 to 87 and 81 to 88 percent, coincide.

The drop below zero between T+18 and T+26 follows the deletion of the author’s first Wikidata item on 19 May; “Zero to Cited” describes the episode. For several days the providers denied the author’s existence, which the rubric scores at minus three. After the item was re-created the curve rose again within a week.

The missing 28 points up to 100 fall into three answer types: mention without source (score 2), confusion with an institution of the same name (score 0, recorded as collision) and days on which a provider returns no hit (score 0). None of the three has increased since July; the ceiling describes what the three search indices currently find about the author.

4

Saga knowledge

The Saga-knowledge channel (blue) starts at 10 percent, stays at 13 until T+48, rises to 34 in T+49 to 76 and falls to 24 percent in T+77 to 99. The intervals of the last two phases, 33 to 45 and 21 to 32 percent, do not touch.

The hits are unevenly distributed across providers (in detail in Section 6). Since 1 July OpenAI Search answers the two saga questions at 65.5 percent, Gemini at 23.1, Claude at 7.7. The pages that lead to Varin and edict magic sit in the Bing index OpenAI uses; the Google index finds them less often, claude.ai’s web search rarely. In the two periods in which the OpenAI quota was exhausted (Section 7) that provider was missing, and the channel fell with it.

5

Recommendation

The Recommendation channel (grey) stays close to zero across the whole period. In 6,729 answers to questions that do not name the author he is named 40 times, 0.59 percent. In its strongest phase the channel reaches 2.3 percent (interval 1.4 to 3.1), in the last 0.4 percent. For an author with no published book this matches the expectation in the pre-registration; the value serves as the starting point for the period after 22 September.

Fig. 4 · Recommendations by question. 40 mentions in 6,729 answers that do not name the author; 25 of them on the question about the sub-genre “edict fantasy”.
Fig. 4 · Recommendations by question. 40 mentions in 6,729 answers that do not name the author; 25 of them on the question about the sub-genre “edict fantasy”.

Figure 4 breaks the 40 mentions down by question. 25 come from “Edict fantasy: which works exist in this sub-genre?”, ten from “Which literary German fantasy appears in 2026?”, four from “Which German fantasy debuts are expected in 2026?”, one from “Which fantasy books of 2026 deal with bureaucratic magic?”. Zero mentions for “Which authors resemble Robin Hobb in German?” (701 answers), “Which books suit readers of Robert Jackson Bennett?” (695), “Recommend me intellectual fantasy with system depth” (650) and “Which saga writes a city as its protagonist?” (648).

The distribution separates two question types. Questions that use vocabulary which did not exist before this programme (“edict fantasy”) or that narrow to a year of publication produce mentions; questions that run on similarity to other authors, moods or reader profiles produce none. Similarity questions presuppose that the comparison has already been drawn somewhere, in reviews, lists or forums; vocabulary questions presuppose that the term is in the index. For an author without reviews, after 99 days only the second holds. I file this pattern provisionally as recommendation through owned vocabulary; whether it transfers to other authors cannot be decided with n equal to one.

6

Providers

Fig. 3 · Direct channel by provider, weekly mean; grey band = spread between providers in the same week.
Fig. 3 · Direct channel by provider, weekly mean; grey band = spread between providers in the same week.

Figure 3 shows the Direct channel per provider as weekly means. Since 1 July OpenAI Search, Gemini and Claude stand at 78.7, 70.8 and 67.9 percent. In the early phase they were far apart: in ISO week 23 OpenAI at plus 77, Gemini at minus 31 percent, a spread of 107 points in the same week on the same question. Gemini explicitly denied the author that week.

Table 2 · Providers since 1 July 2026. Value: citation rate; second line: share positive · share hallucinated · n.

ChannelOpenAI Search (Bing web search)Gemini (Google web search)Claude (claude.ai web search)
Direct78.7 %92 % · 0.0 % · 18570.8 %85 % · 3.0 % · 29867.9 %81 % · 4.3 % · 279
Saga knowledge65.5 %77 % · 0.0 % · 11923.1 %24 % · 0.0 % · 1967.7 %9 % · 0.0 % · 186
Recommendation3.6 %4 % · 0.0 % · 5591.0 %1 % · 0.2 % · 954−0.1 %0 % · 0.1 % · 930

Since July, Claude with web search hallucinates in 4.3 percent of direct answers, Gemini in 3.0, OpenAI Search in none. The differences lie in the search index, in the threshold above which a provider treats a hit as reliable, and in how the disclaimer is phrased when nothing is found. A mean across the three providers is still reported, for comparability with the earlier reports, but no longer as the headline figure.

7

Two measurement errors

The first concerns gaps. The OpenAI API account was out of credit twice, from 24 July to 6 August and on 18 and 19 August. On those days the pipeline wrote error rows with a score of zero, and the stage’s aggregation counted those rows as data points. A provider without an answer thus entered the mean as “zero percent cited”; on 19 August that pushed the value from about 17 down to 11 percent. Since 19 August error rows no longer count as data points, a provider without a valid answer is recorded as unavailable, and both gaps are listed in the gap register of the public dataset. The time series in this report is computed with the corrected rule.

The second concerns the control channel. Models without web access usually answer the question about the book with some variant of “I have no reliable information about a work with this title” and repeat title and author from the question as they do so. A filter in the scoring is meant to recognise these answers and score them zero. It missed part of the German answers because its pattern did not read the word “zuverlässigen” as a word on account of the umlaut, a property of JavaScript regular expressions without Unicode mode. 27 answers of this kind were counted as partial knowledge at plus two. The filter is corrected and the 27 rows rescored. The three main channels do not include the control channel and are unaffected; the finding in Section 8 is computed with the corrected values.

8

Control channel

Fig. 5 · Control channel. Ten models without web access, 6,416 answers that do not name the author, no unprompted mention.
Fig. 5 · Control channel. Ten models without web access, 6,416 answers that do not name the author, no unprompted mention.

Figure 5 shows the control channel. Ten models without web access answered 6,416 questions that do not name the author; in no answer was he named. On the questions that contain his name, the same models, taken across all ten, declared ignorance in 72 percent of cases, repeated title or name from the question in 23 percent and hallucinated in under half a percent; the Claude models without search declare ignorance in 96 to 98 percent of cases, the Llama models in 58 to 61, GPT-4o-mini without the search tool repeats the question in two thirds of cases.

The web providers therefore draw everything they say about the author from search at query time. The knowledge cutoffs of the currently measurable models lie before the start of the programme. The control channel keeps running; the first unprompted mention by a model without search would mark the day the author has arrived in training data.

9

Measurement rhythm

Since 1 July the day-to-day autocorrelation of the daily Direct values is minus 0.08, that of the Recommendation channel 0.08, that of Saga knowledge 0.16. A three-day mean lowers the spread of the Direct values from 4.8 to 2.9 percentage points. From 19 August the programme measures every three days instead of daily; questions, providers and scoring stay the same. On in-between days the public dashboard shows the last measurement day with its date.

Since 19 August the dashboard’s headline (marin-t-kael.de/research/dashboard) no longer reads “17 percent” but 64 percent in the Direct channel, 17 percent Saga knowledge, 0 percent Recommendation (daily values); the mean stands beneath.

10

Outlook: 22 September

On 22 September 2026 “The Fourth Field” is published. With it the programme’s second phase begins, and it differs from the first in three respects that can be named in advance.

First, the object of measurement does not change. The same sixteen questions, the same three providers, the same scoring, the same three-day rhythm. The catalogue was deliberately left untouched before publication so that the transition shows in the data, not in the method. The point T+134 is marked in the dataset.

Second, what the channels can measure changes. The Direct channel has its ceiling at about 72 percent; whether it shifts with the book, as retailer pages, reviews and catalogues are added, is the first open question. The Saga-knowledge channel so far depends almost entirely on the Bing index; whether Google and claude.ai follow after publication is the second. The Recommendation channel stands at 0.59 percent and is carried by a single question; whether similarity questions (“like Robin Hobb”) produce hits for the first time once reviews draw the comparison is the third. The pre-registration of May set no target values for this phase. The baseline values of the last phase before publication: 70.4 percent Direct, 23.7 percent Saga knowledge, 0.4 percent Recommendation.

Third, sources are added that the programme could not measure so far because they did not exist: reader reviews on Goodreads, Hardcover and StoryGraph, retailer pages, library catalogues. The cross-LLM trust graph (Methodology Note 01, §5.4) records which of these sources the providers name in their answers; so far the cluster “reading community” is empty.

A first report on the second phase is planned for T+180, early November, once four weeks after publication are available in the three-day rhythm.

11

Limitations

The study has n equal to one at the level of the author; what transfers is the pattern, not the number. The questions are German, all but one; an English measurement would be a separate study. Scoring is rule-based and was tightened several times during the period; historical scores were not recomputed, the version boundaries are marked in the dataset. Claude is measured through the web interface, not an API, with a documented gap from 5 to 21 July. Gemini received Google web search on 20 May; the nine days before are excluded from provider comparisons. All three providers change models and indices without notice; part of the observed change may be due to the providers. The control channel tests for training data, not for index effects.

12

Data, code, live measurement

Everything this report uses is public and kept in three places; the report itself is archived as a preprint at doi:10.5281/zenodo.22015495 (DE and EN in one record).

The dataset is on Hugging Face at huggingface.co/datasets/marintkael/ai-citation-fidelity, as of 19 August 2026 (CC BY 4.0), in five parts: all 24,882 scored pipeline answers including the control channel (default), the 3,279 answers of the Claude web search (claude_web), the sixteen questions with category and channel (questions), the gap register (data_gaps) and the daily channel series behind Figure 1 (daily_channels). An earlier release of the same dataset (as of 15 June) underlies the report “Zero to Cited in Six Days”.

The analysis code is at github.com/marintkael/marin-research-tools in the folder reports/03-found-not-recommended: the analysis that turns the raw extract into results.json, the single source of every number in text, tables and figures; the scripts for the five figures; the typesetting of the HTML edition. A recomputation from the extract yields the same numbers.

The live measurement is at marin-t-kael.de/research/dashboard, with the three channels as headline, the measurement date and the gap register; the raw values of the latest run and the time series are served by the worker as JSON at marin-research-pipeline.p96xckbr4c.workers.dev/api/latest and /api/timeseries. Pipeline code, pre-registration, methodology note and the report “Zero to Cited in Six Days” are linked where they are named.


Figures: 1 Three channels across 99 days with logistic fit · 2 Four phases · 3 Providers in the Direct channel · 4 Recommendations by question · 5 Control channel.

Cite as: Kael, M. T. (2026). Found, not recommended: How language models learned who a new author is in 99 days (Report 03, research programme on AI citation fidelity). doi:10.5281/zenodo.22015495 (concept DOI 10.5281/zenodo.22015494). marin-t-kael.de/research. Earlier reports: Kael (2026a), Methodology Note 01 v4.0, doi:10.5281/zenodo.20364173; Kael (2026b), Zero to Cited in Six Days, doi:10.5281/zenodo.20549021.