AI Engineering
AARIS: on-device review of academic manuscripts
A local-first multi-agent manuscript reviewer that runs on-device by default and returns line-referenced critiques across methodology, literature, clarity, and ethics.
Problem
Peer review is slow and uneven. A manuscript can wait weeks for feedback, and the feedback that comes back is often high-level prose that points at a whole section, so authors are left guessing what to change.
AARIS addresses that. It takes an uploaded manuscript and returns a structured report where every finding is tied to a specific line, quotes the exact text, names the issue, and proposes a fix. It gives authors and editors a quick first pass specific enough to act on; it does not replace an editor.
- Manuscripts are reviewed on-device by default, so unpublished work never has to leave the machine.
- Every finding in a review is anchored to a line.
- Four independent review lenses run on the same manuscript so no single concern dominates.
- The manuscript domain is detected first, so a medical paper and a mathematics paper are not judged by the same checklist.
How it works
The backend is a FastAPI service. A submission is parsed from PDF or DOCX into plain text, then handed to a LangGraph StateGraph that runs four nodes in order: initialize, create_embeddings, parallel_reviews, synthesize. State moves through the graph as a typed dictionary, and every node writes a checkpoint so an interrupted review picks up where it stopped.
The initialize node detects the manuscript domain and loads domain-specific prompt guidance and scoring weights. The create_embeddings node chunks the text and stores vectors in MongoDB Atlas Vector Search for retrieval during review. The parallel_reviews node fans out to four specialist agents at once, gathered under a timeout. The synthesize node assembles their findings into one editor-style report.
Each specialist agent numbers the manuscript by line, sends a formatted prompt to the language model, and parses the reply as JSON with a strict finding shape. Findings without a line reference are repaired before they reach synthesis, and a re-try wrapper re-asks the model when the reply misses the required structure. Inference is local-first: the default provider is Ollama running on the machine (mistral by default, llama3.2 as the local fallback), so manuscripts do not leave the host. Cloud providers (Groq, OpenAI, Gemini, Claude) sit behind the same per-provider circuit-breaker failover for when a local model is unavailable, with responses cached and an SSRF allowlist on the provider endpoint.
Upload (PDF/DOCX) -> DocumentParser -> Orchestrator
-> LangGraph[ initialize (domain detect) -> create_embeddings (Atlas vector search)
-> parallel_reviews { Methodology | Literature | Clarity | Ethics } -> synthesize ]
-> Report (markdown + PDF) -> MongoDBHard parts
- Format discipline: line-numbered input, a required finding schema, format validation, and bounded re-tries keep findings anchored to real lines; synthesis copies agent findings verbatim so line numbers are never paraphrased away.
- Async/sync boundary: LangChain runs Atlas similarity search synchronously, so the vector store uses a dedicated PyMongo client while the rest of the app stays on async Motor, avoiding an async-cursor mismatch.
- Failure isolation: parallel reviews run under a gather-with-timeout, per-agent exceptions are caught and recorded, and a failed lens degrades to a placeholder and the rest of the review still completes.
- Resumability: checkpoints at each workflow stage plus a stuck-submission recovery path let a review continue after an interruption.
- Report resilience: if synthesis fails or drops the line-by-line format, an emergency report is built directly from the collected agent findings.
Results
What the system does is easier to pin down than how well it scores; the repository does not define benchmark numbers, so the outcomes below are behavioural.
- Four review dimensions produced per manuscript, each carrying up to roughly a dozen line-referenced findings.
- Domain detection covers a broad set of academic disciplines, with per-domain prompt guidance and scoring weights.
- Local Ollama inference by default so manuscripts stay on the machine, with cloud providers behind circuit-breaker failover and an ethics path that can seek multi-model consensus.
- Reviews delivered as both a structured markdown report and a generated PDF.
- The workflow is checkpointed at every stage and resumes after an interruption.
Artifacts
The repository is private, so no source link is included.
- FastAPI backend with versioned routes for submissions, author/editor/admin dashboards, authentication, and system health.
- LangGraph review workflow and four specialist agent implementations.
- RAG layer over MongoDB Atlas Vector Search with an embedding cache and a RAG metrics endpoint.
- Security layer: dual JWT and API-key auth, role-based access control, WAF, rate limiting, Keycloak and Vault integration, field encryption, and audit logging.
- Operational tooling: checkpoint and stuck-submission recovery, dead-letter queue, cost monitoring, data-retention and GDPR export, and PDF report generation.