Ask the Stanford Encyclopedia of Philosophy
Live answers may take a minute or longer; one question runs at a time.
Ask a philosophy question and inspect the encyclopedia passages behind the answer. Ask the SEP searches the Stanford Encyclopedia of Philosophy, drafts an answer with citations, and shows a model’s review of whether those citations support the claims. I studied philosophy as an undergraduate, which is why this is the second corpus: it is the one where I can tell when an answer is subtly wrong.
In a businessA document assistant that works on one archive can stumble on another archive's vocabulary. This project moves the Ask Mindscape workflow to an encyclopedia and keeps its evaluation separate, since results on the podcast do not carry over.
Technical skills: Python · FastAPI · LangGraph · MCP · RAG · Corpus-specific prompt design
- My contribution
- I adapted the Ask Mindscape workflow to the encyclopedia and reviewed its prompts. A corpus profile replaces the prompts and page copy; the workflow, the review, and the separation between drafting and revealing evidence are shared with Ask Mindscape.
- Result and limit
- Live and not yet evaluated. The Mindscape source-recovery measurements have no SEP counterpart; a four-question development pilot is not evidence of answer quality.
On this page
A useful starting question is “What is a Gettier case, and why does it show that justified true belief is not sufficient for knowledge?”
Technical methods and evaluation
Jump to encyclopedia-specific prompts, evaluation limits, or the implementation.
What changes for an encyclopedia
The prompts differ because the material differs. A transcript has speakers, so the Mindscape prompts separate what Sean asks from what he endorses and what a guest claims. An encyclopedia entry may describe positions, objections, and replies without endorsing each of them. A draft can therefore misattribute a claim if it presents “what Lewis holds” as an uncontested fact. The SEP prompts ask the draft to name whose position each claim is, to keep to the question asked rather than surveying related material, and to say plainly what the passages do not cover.
The review stage is asked a stricter question than topical relevance: does each claim follow from the passage it cites? A passage on the same subject does not necessarily support the claim. Attribution problems it looks for include a position an entry describes presented as the entry’s own view, an objection presented as the position it targets, and a philosopher credited with a view the passage does not give them.
Passages are short quotations linked to their entry and section. The reranking budget is the same 12 s as the other pages; the SEP instance has four times the chunks of Mindscape, so the floor is more often labeled partial.
Evaluation and limits
This version has not been evaluated. The Mindscape measurements rely on listener questions whose answers are on record. An encyclopedia evaluation needs a separate set of reference questions and answers. A four-question pilot run during development found the pipeline runs and showed the kinds of attribution mistakes the prompts above are written against; it is not evidence of answer quality. Treat this as a demonstration of the workflow on a second corpus, with the same separation between drafting and revealing the connection evidence.
The extra passage is a probe, and it can be off-topic. Its connection comes from a shared feature of the text, so a question about inherited claims to reparations can bring back an entry on biological inheritance. Nothing on this page measures whether it helps. On the podcast archive, answers cited it in 2 of 46 questions where one was seated.
The model is the visitor’s choice: a locally served Qwen checkpoint, or a hosted GLM or DeepSeek model, which is faster and streams its draft as it writes. Answers can be wrong, and the review can be wrong about the draft; the passages are there so the reader can check.
My contribution and stack
I adapted the Ask Mindscape workflow to the encyclopedia and reviewed its prompts. One setting points the Ask Mindscape workflow at Heart of Gold’s encyclopedia index: 1,795 entries in a frozen snapshot, split into 85,263 passages. Retrieval, drafting, model review, answer hashing, and evidence reveal are shared. The corpus profile replaces the four blind-stage prompts and the page copy; everything else is shared with Ask Mindscape.
Python, FastAPI, LangGraph, Heart of Gold over MCP, a locally served Qwen model or a hosted one (GLM, DeepSeek) chosen per question, and server-sent events. The demo is served through a Cloudflare tunnel.
Where this could be useful
The same workflow could answer questions over a reference collection whose documents describe competing positions — legal commentary, standards, policy reviews — where the reader needs to know whose view a sentence reports. A pilot would need representative questions and an independent check of attributions before being used for decisions.
If you are hiring, the experience page says which business problem this project shares its mechanics with.