A retrieval system that is allowed to surprise you
On this page
A retrieval system is usually evaluated on whether it finds material relevant to a question. That is a reasonable priority, but it leaves open a different question. Could the system also help someone encounter a useful connection they would not have thought to search for? Similarity-based retrieval can produce those connections, but it does not specifically distinguish them from more obvious results.
Research on recommender systems treats relevance and unexpectedness together under the concept of serendipity. I wanted to investigate whether a similar approach would be useful for retrieval-augmented generation. Heart of Gold combines an ordinary ranked set of passages with additional passages selected through a statistical test of connections in the corpus. I built it first over the Stanford Encyclopedia of Philosophy and then over the Mindscape transcripts. This post describes the encyclopedia version.
An unexpected connection
I asked about a question from a paper I wrote in 2017. Can the passage of time weaken descendants’ claims to restitution? Alongside passages about reparations, Heart of Gold suggested an entry on biological inheritance systems. Both entries concerned what passes to descendants and on what grounds. I would not have searched for the biology entry.
Two other suggestions failed. One connected Black feminist philosophy to black holes; the other was a name-based artifact. All three had passed the statistical test. It would therefore be a mistake to treat passing that test as equivalent to finding a useful connection. The biology comparison gave me something relevant to consider, while the other two showed that a strong numerical association can arise from a confusion of names or meanings. The full retrieval and my verdicts appear below.
Technical methods and evaluation
Jump to passage retrieval, the connection test, or the full worked example.
The rule the whole thing is built on
The name comes from the ship with the Infinite Improbability Drive in The Hitchhiker’s Guide to the Galaxy. That suggested a design in which the user sets a level and the system samples from an eligible set. I was concerned that optimizing a single score for surprise could reward results that score well without being useful. Heart of Gold therefore applies conditions that determine which connections are eligible, then samples within that set. This avoids always selecting whichever connection maximizes the score. However, a poor connection can still satisfy the conditions, as the black-holes example shows. The sampling rule addresses how connections are selected; whether the eligible connections are useful remains a separate question.
The floor
I first need a useful set of answers to the question itself. Extra connections should have something relevant to connect to.
The ranked passages form what I call the floor. The query is embedded twice, once as a question and once as a document, because the embedding model handles those inputs differently. A cosine search returns a pool of 376 chunks, and a lexical channel adds candidates containing rare query terms. A cross-encoder reranker, Qwen3-Reranker-0.6B, scores each candidate against the query, and the top 28 form the floor. The reranker processes about 32 chunks a second on my machine. Interactive requests have a deadline, and a response is marked partial if that deadline expires before the whole pool is scored. The scores and number of scored candidates are returned with the passages.
The certificate
The connections used for the additional passages are computed offline. A sparse autoencoder represents each chunk through a small set of active features from a total of 17,281. The system calls each feature a marker. If two encyclopedia entries share an active marker, they are candidates for a connection, subject to the following checks.
Two entries might share a feature simply because they discuss much the same subject. If so, the overlap gives little reason to expect an unexpected connection. I therefore compare the feature’s strength with that of pairs whose articles are similarly close overall. This asks whether there is something more specific to account for than their general similarity.
The co-occurrence must be at least 28 times the chance expectation. Its evidence, measured as the negative log tail probability of the joint, must also exceed a high quantile within the pair’s kinship bin. Kinship is the cosine similarity of the two entries as whole documents. Comparing pairs within a kinship bin helps distinguish a connection through a particular feature from the simpler explanation that the entries are generally similar.
Both entries must contain the marker in more than one chunk, and they must not share a quotation of 28 tokens or more. The quotation condition excludes some pairs that are connected only because they quote the same passage. Finally, the entries must be at least three steps apart in the encyclopedia’s related-entries graph. This excludes direct editorial links and links through a single intermediate entry. Together, these conditions remove several ordinary explanations for the overlap. They still leave open whether the remaining feature means the same thing in both entries and whether comparing the entries helps with the question. Those are the issues the worked example below examines.
There is also a more basic objection: perhaps the procedure finds apparently strong connections just as readily in arbitrary pairs. The placebo comparison checks this by applying the same test to matched random pairs. A run fails if too many placebo pairs pass. For example, one version produced 1,704 target connections against a placebo mean of 748 and failed this check. That result is recorded in the notebook.
On the December 2023 snapshot of the encyclopedia, which is 1,795 entries in 85,263 chunks with 19,537 resolved citation edges, the certificate yields 2,693 pairs.
Seating
The final choice needs to avoid filling every extra slot with variations on one connection. The selection therefore spreads the added passages across different features and types of connection.
For a query, an eligible pair must have one entry on the floor and one off it, with its marker active in the query or the floor. The system then fills available slots, called seats, with no more than one passage per marker. Selection is spread across signature classes, which describe the resolutions at which the connection appears. It also gives priority to markers used less often during the session, under a leximin rule. The draw within the eligible groups is uniform.
The dial offers no seats, one seat, or up to 28 seats. If there are too few eligible connections, the system returns fewer passages and reports the shortfall. It does not add unrelated passages to fill the requested number.
The additional passages are mixed into the floor without being identified as special. In the research client, the model must finish its answer before it can see the panel identifying those passages. The answer’s hash is recorded so that a later change can be detected. This separation is intended to prevent the model from writing an answer around a passage simply because the system has called it surprising.
The model panel can help assess the suggestions, but its agreement cannot settle whether a reader finds them useful. The panel is another model-based assessment of the same material, while usefulness is the question for which the system needs human judgments. I therefore record my judgments separately, so that the panel’s labels do not stand in for that evaluation.
The words the panel uses
The panel uses a few terms that are useful to define before reading the example.
| Term | Meaning |
|---|---|
| Marker | A sparse-autoencoder feature shared by a pair of entries. |
| Seat / slate | A slot for an additional passage, and the set of slots filled for a query. |
| Kinship | Similarity between the entries as whole documents. |
| Evidence | Strength of the co-occurrence signal relative to other pairs for that marker. |
| Signature | A four-bit record of whether the connection appears at coarse, mid, fine, and marker resolutions. |
| Leave-support-out (LOO) | The signature recomputed after removing the marker’s supporting passages. A changed bit indicates that those passages affect the connection under that measure. |
| Act-symmetry | How similarly the marker activates on the two sides. A value near 1 indicates similar activation strength, which need not mean the words have the same meaning. |
| Petunia / whale | My labels for a useful, non-obvious connection and an artifact or trivial connection, respectively. |
The last two labels also come from The Hitchhiker’s Guide to the Galaxy. They are human verdicts, distinct from the statistical certificate.
What happened when I asked it a question
I used a question from my 2017 paper on territorial rights. Can a right be superseded over time so that descendants of people who were wronged lose a claim to restitution? The query used dial 3 and displayed eight of the 28 floor passages. The panel below is reproduced as the server returned it.
The floor included passages from territorial-rights, justice-intergenerational, and black-reparations. These covered Waldron’s waterhole case, the distinction between restitution and reparations, and the question of whether a right grounded in particular interests changes when those interests no longer hold. These were relevant to the argument I had made in the paper.
seating: 28 seat(s) offered, 3 filled, one per marker; eligible pool 3 pair(s) over 3 marker(s)
pool built from 9 certified pair(s) touching the 7 floor entries: 0 both ends on the floor,
6 marker outside the query/floor feature profile, 0 author-ruled whale(s)
undersubscribed: only 3 distinct marker(s) were eligible
black-reparations <-> inheritance-systems
via inheritance · you · fly · bats (f9572)
evidence_rel 19.95 | chance x131.5 | kinship 0.369 (pctile 20.0) | act_symmetry 0.669
signature (1,1,1,1) (LOO (1,1,1,1)) | graph distance 4 | shared citations 2
black-reparations <-> spacetime-singularities
via Black feminist philosophy (f4379)
evidence_rel 23.02 | chance x82.6 | kinship 0.351 (pctile 15.1) | act_symmetry 0.908
signature (1,1,1,1) (LOO (1,1,1,1)) | graph distance 4 | shared citations 2
black-reparations <-> edith-stein
via Susan Stebbing (f6445)
evidence_rel 28.11 | chance x56.6 | kinship 0.37 (pctile 20.3) | act_symmetry 0.651
signature (1,1,1,1) (LOO (1,1,1,1)) | graph distance 3 | shared citations 1
Three seats filled of 28 offered, because only nine certified pairs touched the floor and six of those had markers outside the query’s profile. My verdict, recorded in the author column, is two whales and a petunia.
I judged the spacetime-singularities connection to be a whale. The marker’s top tokens concern Black feminist philosophy, but it also activates on the singularities entry because of black holes. The act-symmetry score was 0.908, so similar activation strength did not distinguish these different meanings. The edith-stein connection was another whale, produced through a marker associated with Susan Stebbing. In the first case, the numerical similarity is compatible with a clear difference in meaning. The annotations describe how the feature behaves, but the objection concerns what it refers to. A strong score on the former cannot answer the latter.
I found the inheritance-systems connection useful. The reparations entry discusses an inheritance argument under which descendants of enslaved people inherit a claim as heirs. The biology entry concerns the mechanisms through which characteristics pass between generations. These are different kinds of inheritance, but comparing them raised a question relevant to the paper. What is being passed to descendants, and what grounds the claim that it should pass?
The entries are four steps apart in the editors’ graph, share two citations, and have whole-document similarity in the twentieth percentile. Those measurements explain why the pair was eligible; my judgment that the comparison was useful goes beyond them. The query also brought me back to the seminar in which I wrote the paper and read Tommie Shelby’s Dark Ghettos. I would not have thought to search for the biology entry myself.
Where it stands
The worked example establishes that the system can return a connection I find useful alongside suggestions I reject. How often it does so is a different question. At the time of this example, I had recorded five author verdicts, compared with 2,693 certified pairs and 2,919 rows labeled by the language-model panel. A set of 129 pairs was reserved for human calibration, but I had judged only one of them. With so little of that set assessed, the example cannot support an estimate of human precision. Completing that evaluation is what would let me move from describing an interesting case to making a broader claim about the system’s quality.
The same retrieval engine also serves the Mindscape archive, where passages include timestamps and speaker information. Its floor powers the Explorer’s search (opens in a new tab). Both instances use an MCP server with model-readiness reporting, time budgets, labeled fallbacks, and a log of each query and the floor returned. The project page collects the operational measurements.