Scientific Research Assistant
A command-line research system that treats a study as explicit, versioned state rather than a conversation: eleven specialist capabilities handle research questions, methodology, evidence, statistics, ethics, review and document generation, and every consequential field changes only through a recorded, approved decision.

The problem
Research is a chain of dependent decisions. A topic becomes a question, the question implies a design, the design fixes the variables and the analysis, and the analysis has to be defensible for that method. Generative tools tend to jump to prose, which produces a polished document resting on a methodology nobody agreed to.
Traceability is the second failure. A reference that cannot be verified is indistinguishable from a real one, an assumption made during drafting looks like an established fact, and a method changed in one chapter is not reflected in another.
System objective
Model the research lifecycle as application state rather than as a prompt, so that quality follows from structure instead of from how a request was phrased — and so a researcher can see, approve and reverse every consequential decision.
- Never fabricate. Where a value cannot be determined, the system writes a [TBD] placeholder or records it as not available rather than inventing a plausible one.
- Consequential fields change only through an approved decision, enforced in code rather than by instruction.
- Every state change is a committed version with a timestamp, an actor and a reason.
- Degrade honestly — an unavailable literature feed or reasoning model is labelled as such, never silently substituted.
Who it is for
Built for the researcher carrying a structured obligation — an undergraduate or master's project, a doctoral thesis, a clinical or professional study, or a grant or funder proposal.
- Framing a researchable question from a broad topic using an appropriate question framework.
- Selecting and justifying a study design, population, sampling approach and sample size.
- Building an evidence base from verified references rather than plausible-looking ones.
- Deciding what the ethics position is, and recording it without ever self-granting approval.
- Drafting and exporting the resulting proposal as Markdown, Word or PDF once the quality gate passes.
System architecture
The system is pure Python 3.10+ using the standard library, with optional extras for richer document output. A command-line entry point creates a project, opens an interactive session on it, routes a plain-language request to the right capability, and writes the resulting state change as a version.
Routing is deterministic and ordered: meta commands first, then a capture check, then the longest matching keyword trigger, defaulting to research-question framing. Eleven capabilities sit behind that router — capture, question, methodology, evidence, statistics, ethics, writing, citation, review, validation and documents — and each declares which fields it may write. Reasoning, literature, document rendering and statistics sit behind swappable provider interfaces.
- Command-line interface — new, list, open, ask, decide, docs and test. Each project is a plain folder that can be opened, copied or backed up. There is no database and no server.
- Capability registry — eleven capabilities registered behind one router, each declaring the fields it is allowed to set, so authority is structural rather than conventional.
- Approval gate — consequential fields (research question, aim, objectives, study design, primary outcome, sample size) are proposed, then applied only when a decision code is approved. Rejecting leaves the previous value intact.
- Versioned project state — every change is committed as an immutable version snapshot, with an append-only audit trail recording timestamp, actor and reason, and any two versions can be compared.
- Evidence with verification — literature search reaches PubMed, Europe PMC and Crossref, and a candidate identifier is verified against those sources before a reference may be treated as authoritative.
- Deterministic statistics — sample-size and power calculations are computed rather than estimated, and stored as declared assumptions instead of appearing as unexplained numbers in a draft.
- Honest failure modes — unavailable reasoning returns a deterministic provisional question with placeholders, an offline literature feed yields not available rather than a plausible citation, and a missing document formatter is reported as skipped.
- Swap-on providers — reasoning, literature, document and statistics providers are behind interfaces, so the system is not bound to one model or one data source.
Key features
- Study as versioned stateThe project holds an explicit record — question, design, population, variables, outcomes, sampling, analysis plan, evidence, assumptions, limitations — with immutable version snapshots and an append-only audit trail, rather than relying on conversation history.
- Approval gate on consequential fieldsStudy design, aim, objectives, primary outcome, sample size and variables are proposed and explained, then applied only through an approved decision. The gate is enforced in the state layer, so no code path can write around it.
- Anti-fabrication as a design ruleUnverifiable citations are marked unverified, and the system states not available rather than producing a plausible reference. Ethics approval is recorded as a note and approval status stays not obtained, because the assistant does not grant it.
- Verified evidence retrievalLiterature search against PubMed, Europe PMC and Crossref, with identifiers verified against those sources before a reference is treated as authoritative.
- Deterministic statisticsSample-size and power calculations are computed rather than estimated and stored as declared assumptions, so the basis of a number in the document is inspectable.
- Quality gate before documentsDocuments are only written when no critical review finding is open and the protocol is coherent. Producing them anyway is possible but is recorded on the document as forced.
How it was built
The central decision was to make integrity a property of the software rather than advice in a prompt, because a prompt-level instruction is exactly what fails when the user is in a hurry.
- Standard library first — the core runs on Python's standard library alone, so the system has no install-time dependency and cannot break because a transitive package moved.
- Authority declared in code — each capability states the fields it may write, and the state layer independently refuses a consequential write that has no approved decision behind it.
- Routing order is explicit and tested — meta commands, then capture, then longest keyword match, so behaviour does not change when a new trigger phrase is added.
- Providers behind interfaces — reasoning, literature, documents and statistics are swappable, and each degrades to a labelled, honest fallback instead of a guessed answer.
- Hermetic tests — the suite covers state, decisions, capture, routing, review, statistics, validation, workflows and safety, and runs offline so a network failure cannot mask a logic failure.
- The researcher's own wording is preserved — a variable the researcher states explicitly is applied as stated; the system does not rewrite it.
Design decisions worth noting
Several decisions trade convenience for a property the domain actually requires.
- No database and no website — a project is a folder of files, which means it can be opened, copied, diffed and backed up. There is no service to be unavailable.
- Documents are gated, not merely warned about — the export path refuses to run while critical findings are open, so an indefensible document cannot be produced by ignoring a message.
- The system will not grant ethics approval — it records the ethics position and keeps approval status at not obtained, because that is a decision for an ethics committee and a human.
- Offline is a supported mode — the system is designed to be usable without a reasoning model or a literature feed, degrading to placeholders rather than inventing content.
How a study progresses
- 01 — Create the projectLevel, mode, discipline, context
- 02 — Open a sessionInteractive or one-shot request
- 03 — Route the requestLongest-match capability trigger
- 04 — Propose a consequential fieldExplanation, never a silent write
- 05 — Approve the decisiondecide <code> approve
- 06 — Commit a versionTimestamp, actor, reason
- 07 — Retrieve and verify evidencePubMed, Europe PMC, Crossref
- 08 — Compute the analysis planSample size and power
- 09 — Pass the quality gateNo open critical findings
- 10 — Export the documentsMarkdown, Word, PDF
Current status
The system is working and covered by a test suite that runs offline. All eleven capabilities are implemented behind the router, with the approval gate, versioning, audit trail, evidence verification, deterministic statistics and document export in place.
The document renderers are optional extras: with python-docx and reportlab installed it produces real Word and PDF output, and without them it degrades to Markdown and text rather than failing.
This is the broader research system. Its browser-based implementation is documented separately as Scientific Research Assistant — Web Edition.
Delivery status
- CompleteProject folder, manifest & state model
- CompleteCapability router with ordered triggers
- Complete11 specialist capabilities
- CompleteApproval gate on consequential fields
- CompleteImmutable versions & audit trail
- CompleteEvidence retrieval & reference verification
- CompleteDeterministic sample-size & power engine
- CompleteEthics handling without self-approval
- CompleteCritical review & quality gate
- CompleteDocument export (Markdown, Word, PDF)
- CompleteOffline hermetic test suite