Maxwell Grody

Ask Mindscape

Answers across podcast transcripts · inspectable AI workflow

Ask the archive ↗ (opens in a new tab) · Ask the SEP ↗

Live answers may take a minute or longer. A recorded example is available in the demo.

A question can span many episodes of Sean Carroll’s Mindscape podcast. Ask Mindscape retrieves transcript passages and uses a language model to draft an answer across them. The public interface shows the stages of that process, so you can inspect the source material, the draft, the verification notes, and the final answer.

In a businessA team searching its own policies or records needs answers it can check against the source. This project evaluates finding the right source and writing the answer separately, because a found source does not make the answer right.

Technical skills: Python · FastAPI · LangGraph · MCP · Retrieval-augmented generation (RAG) · Source-recovery evaluation

My contribution
I set the project requirements and built the service. The interface exposes retrieval, drafting, verification, and an answer saved before the model learns which connection was added.
Result and limit
Across 50 archive questions, the known answer location was retrieved in 35 cases and cited in 34. This measures source recovery, not answer accuracy; median runtime was 47 seconds.
On this page

The live index now covers 441 episodes through 2026-09-21, refreshed on 2026-09-27. The evaluation below remains a measurement of the original 436-episode snapshot, not a new test of the expanded archive. Existing additional connections retain their original model and certification snapshot; new episodes have been added to relevance search.

A useful starting question is “What different meanings of emergence have appeared on Mindscape?”

Technical methods and evaluation

Jump to the answering steps, evaluation, or runtime.

What happens to a question

  1. Retrieve. Heart of Gold finds relevant passages. With the connection setting enabled, it can also include an additional passage selected through the retrieval system’s connection evidence.
  2. Draft. The model sees the passages without being told which one is the additional connection. It writes an answer with numbered citations.
  3. Verify. A model-based review checks the draft against the passages. The interface displays the revised answer and review notes.
  4. Freeze. The answer is fixed and hashed before the connection evidence is revealed.
  5. Reveal. The interface identifies the additional passage, if one was available, and reports whether the answer cited it.

The stages stream to the browser as they finish. Readers can follow citations back to the displayed passages and their episode sources.

Why hide the connection label?

If the drafting model is told which passage is the special connection, it may favor that passage because of the instruction. Hiding the label makes it possible to observe whether the passage gets used without that explicit cue.

The graph separates the drafting tools from the evidence-reveal tools. Reveal happens after the answer is frozen. Tests check the event order and that a hidden evidence marker does not appear in the draft or verified answer.

This controls one source of bias. It does not establish that the connection improved the answer, or that the model cannot infer which passage is unusual from its content. A hash records the frozen answer; it is not a certificate of factual correctness.

How this fits with the other projects

These projects share infrastructure and source material. Ask Mindscape adds the orchestration, verification, and streaming answer interface.

Evaluation and limits

The evaluation asks 50 listener questions sampled from the AMA question–answer index (seed 20260910), using Qwen3.6-35B-A3B at 8-bit with the connection setting enabled (dial 2). The reference for each question is the episode and time span of Sean’s recorded answer; the answering workflow never sees those labels. All 50 runs completed. The figures below are from the second run of this evaluation, on the same 50 questions as the first; what changed between the runs is described after the table.

A retrieved passage counts as a source match when it comes from the same episode and its chunk’s time span overlaps the recorded answer span, extended by 90 seconds on either side. A chunk runs from its start timestamp to the next chunk’s start, about two minutes of speech. This is a test of recovering a known answer location. It does not establish that a passage supports a generated claim or that the generated answer is correct, and these are existing archive questions, not unseen topics.

The 95% interval describes uncertainty around each measured success rate using the Wilson method. Wider intervals reflect the limited number of test questions; they do not capture every way new questions might differ from this sample.

Measure Result 95% interval
Known answer location among retrieved passages 35/50 · 70% 56–81%
Same, by the first run’s rule (chunk start inside the span) 24/50 · 48% 35–61%
Recorded answer present in the text the model saw 35/50 · 70% 56–81%
Known answer location cited in the frozen answer 34/50 · 68% 54–79%
Cited, when retrieved 34/35 · 97% 85–99%
Additional connection passage included 46/50 · 92% 81–97%
Additional passage cited, when included 2/46 · 4% 1–15%

What changed between the two runs

The first run reported the known answer location among the retrieved passages in 25 of 50 cases and cited in 23. An audit of three of its misses found that two were not retrieval failures. In both, retrieval had selected the chunk containing Sean’s answer, but only the chunk’s first 120 words reached the model, and those words were the previous listener’s question. The scoring rule compounded the problem: it accepted a chunk only if its start fell within 90 seconds of the answer, and a chunk that begins two minutes earlier and contains the whole answer failed that test.

Two changes followed. The passage each model sees is now a 120-word window placed where the question’s words fall in the chunk, rather than at the chunk’s opening. The window depends only on the question and the chunk text, so it is the same for the ordinary passages and the additional one. And a chunk now counts as a source match when its span overlaps the answer, with the earlier rule reported alongside for comparison. Retrieval itself did not change: the two runs agree question by question under the first run’s rule on 49 of 50, and the one disagreement is a partially reranked pool landing differently. The revised matching rule accounts for the higher source-recovery count. The citation increase is consistent with the revised snippets exposing more relevant text, but this comparison changes both snippet selection and scoring; it does not isolate their individual effects.

The third row is a text check rather than a timestamp check: whether the words the model saw share a long run with the recorded question and answer, beyond what the same passage shares with any other sampled question’s text. It agrees with the span rule on all 50 questions. So in the 15 misses, no retrieved passage carried the recorded answer’s wording from another episode. Mindscape revisits subjects across interviews, solo episodes, and AMAs, and a different episode may still offer adequate evidence in other words, which a text check cannot see; that remains a plausible explanation for some misses rather than a finding.

The additional passage was rarely cited in either run. This provides no evidence that adding it improves answer quality: there is no paired connection-off comparison or independent human assessment here, and citation frequency alone would not establish usefulness even if it were higher.

What the verifier reported

The verifier emitted at least one unsupported-claim note for 46 of 50 answers and at least one attribution note for 44, striking 1.14 claims per answer on average while shortening answers by about 1% (3,054 to 3,015 characters on average). These are the model’s own flags, not confirmed error or repair rates; in the first run some attribution notes described an attribution as correct. I would not read these counts as “92% of answers were corrected.” Independent review of paired drafts and revisions is needed to establish whether verification helps.

Latency and partial reranking

Stage Median Mean
Retrieval 13 s 13 s
Drafting 13 s 13 s
Verification 17 s 53 s
Full run 47 s 80 s

Stage medians do not add up to the median full-run time. Two runs took over 300 seconds (769 and 773 s), both in verification, and they account for most of the gap between typical and average latency. The 4,096-token cap on any one generation turns a runaway review into a failed validation and one repair round. It does not shorten the generation itself. These timings describe this evaluation, not a response-time guarantee for the live service.

All 50 runs label reranking as partial: the 12-second budget scored part of the candidate pool (the first run’s generated report listed zero because it read an empty note field; the report now reads the label itself). The retrieval results are results for the budget-limited configuration.

Download the aggregate evaluation results. The download records counts, intervals, timings, and a hash of the source evaluation file, without the generated answer texts. Each evaluation row keeps every passage as the model saw it, with its chunk and time span, so citations can be checked later.

The service also has tests for citation and timestamp parsing, chunk-span scoring, the passage window, streamed event order, freeze-before-reveal behavior, and failure/readiness handling. Those tests check implementation behavior rather than answer quality.

Questions, timings, and run diagnostics are logged by the demo for analysis; they are separate from the main website’s aggregate traffic counters.

The same workflow on the encyclopedia

Ask the Stanford Encyclopedia of Philosophy points this service at Heart of Gold’s SEP instance. One setting selects the corpus; the prompts change because an encyclopedia reports positions without endorsing them. That version does not yet have a comparable evaluation. It needs its own reference questions and answers; the podcast’s recorded AMA answers cannot serve as references for the encyclopedia.

My contribution and stack

I set the project requirements and built the service. The interface shows the passages, draft, and review notes so readers can check how the answer was constructed. I also required the answer to be fixed before the model learns which passage was added as an unexpected connection.

Python, FastAPI, LangGraph, Heart of Gold over MCP, a locally served Qwen model or a hosted one (GLM, DeepSeek) chosen per question, and server-sent events. The demo is served through a Cloudflare tunnel.

Where this could be useful

A similar workflow could support research across an internal document collection, where readers need both an answer and the evidence behind it. A pilot on a team’s own documents would need its own access controls, representative questions, and checks for unsupported answers before being used for decisions.

If you are hiring, the experience page says which business problem this project shares its mechanics with.