A Streamlit-based application that analyzes academic papers using Retrieval-Augmented Generation (FAISS + semantic search). Provides cited answers, evidence extraction, and instant summaries of research papers.
This is a Streamlit frontend application, not an MCP server. The codebase contains 8 functions exposed as tools, but they lack formal MCP tool registration, schemas are partially inferred from docstrings, and descriptions are inconsistent. Most critically: no evidence of MCP protocol implementation (no `ListTools`, `CallTool`, `Initialize` handlers, no JSON-RPC transport). The code is a Streamlit chat UI that directly invokes Python functions, not a remote tool provider. Tool definitions are scattered across function docstrings with minimal structure. Several tools accept untyped or loosely-typed parameters (e.g., `pdf_file` is documented as 'object' from Streamlit uploader, not a formal schema). Error handling is absent in most tools, no try-catch blocks, no structured error responses, no recovery guidance. Output schemas are undocumented. This scores low because it is fundamentally not an MCP server and lacks the formality required for agent integration.
Builds a new FAISS vector store from text or loads an existing one from disk. Uses HuggingFace embeddings with sentence-transformers/all-MiniLM-L6-v2 model.
ensure_chunksread onlysource verified53/100
Splits text into chunks using RecursiveCharacterTextSplitter with chunk_size=800 and chunk_overlap=150, then filters out noisy chunks containing links or being too short.
free_answerread onlysource verified60/100
Returns (answer, score). Always returns 2 values.
get_index_folderread onlysource verified55/100
Generates a safe folder name from a file name by keeping only ASCII letters/numbers/_/- characters.
get_qa_modelread onlysource verified48/100
Cached resource that loads and returns a question-answering model (deepset/roberta-base-squad2).
Not an MCP server, this is a Streamlit frontend application with no MCP protocol implementation. No evidence of Initialize/ListTools/CallTool handlers, no JSON-RPC transport, no client-server separation.
Tool definitions are Streamlit-coupled and not portable. E.g., `read_pdf_text` accepts a Streamlit file uploader object; `build_or_load_vectorstore` uses `st.session_state` for caching; `get_qa_model` uses `@st.cache_resource`. These are framework-specific and cannot be used by an MCP client.
No input schemas provided in formal MCP format. Parameter types are mentioned in docstrings (e.g., 'question:string') but not registered as JSON Schema. Schemas are inferred from function signatures, not explicitly declared.
Recommendations
Refactor this Streamlit app into a proper MCP server. Implement an MCP-compliant transport (HTTP or STDIO) and register tools using the MCP Protocol's ListTools and CallTool handlers. See https://modelcontextprotocol.io/introduction for specification.
Define formal JSON Schema for all tool inputs and outputs. Example for `free_answer`: {"type": "object", "properties": {"question": {"type": "string", "minLength": 5, "maxLength": 500, "description": "The research question to answer (5-500 chars)."}, "context": {"type": "string", "minLength": 50, "maxLength": 10000, "description": "The paper context to search within (50-10000 chars)."}}}}. Declare output schemas: `free_answer` returns {"type": "object", "properties": {"answer": {"type": "string"}, "confidence": {"type": "number", "minimum": 0, "maximum": 1}}}.
Expand tool descriptions to 100-200 characters, following the format: '[Action]. [When to use]. [Prerequisites]. [Returns].' Example: 'Extract a cited answer from the paper context using a lightweight QA model. Use when OpenAI API is unavailable or for fast, offline responses. Requires context >= 50 chars. Returns (answer_string, confidence_0_to_1).'
Add error handling to all tools. Wrap external calls (OpenAI API, PDF parsing, vectorstore I/O) in try-except blocks. Return structured errors with actionable recovery guidance. Example for `openai_answer_once`: {"error": "API rate limit exceeded", "retry_after_seconds": 60, "suggestion": "Wait 60 seconds and retry, or use free_answer() as a fallback."}
Score history
Overall score trend
↑ 16 points across a rubric change (v1 → v2)
40/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
40
2026-07-28+
v2
2026-03-09
F
24
-
v1
auth
source verified
58/100
Streaming answer for chat (OpenAI).
read_pdf_textread onlysource verified53/100
Extracts text from a PDF file by reading all pages and joining them with newlines.
Output schemas are entirely undocumented. No tool declares what it returns (type, structure, fields). Callers must reverse-engineer from code (e.g., `free_answer` returns a 2-tuple; `openai_answer_stream` returns a generator).
No error handling. None of the 8 tools have try-catch blocks or return structured error responses. If a tool fails (e.g., PDF is corrupted, OpenAI API is down, vectorstore is corrupted), the exception propagates unhandled. An LLM has no guidance on what went wrong or how to recover.
Non-idempotent operations without safeguards. `build_or_load_vectorstore` writes to disk (deletes and recreates faiss_indexes/<name>/) but has no dry-run, confirmation, or recovery mechanism. If called twice with conflicting text, the second call silently overwrites the first.
Tool composition is poor. `get_index_folder` and `ensure_chunks` are utility functions with no independent value, they are tightly coupled to `build_or_load_vectorstore` and should not be exposed as separate tools. `build_or_load_vectorstore` combines two concerns (build vs. load) and should be split.
Descriptions are incomplete or vague. Many descriptions (e.g., 'Streaming answer for chat (OpenAI).', 'Non-stream call (useful for summary).') are under 50 characters and lack guidance on when to use each tool or what fields the LLM should expect in the response.
Secrets exposure risk. The code uses `os.getenv('OPENAI_API_KEY')` which, while safer than hardcoding, is still not ideal for an agent-callable tool. If this tool were exposed via MCP, the API key must never appear in tool parameters or responses. Current design is acceptable for a Streamlit app but would need hardening for an agent.
No pagination or result limits. `build_or_load_vectorstore` can return thousands of chunks without pagination. `free_answer` and `openai_answer_*` tools truncate context to 6000 chars (for free_answer) or accept unlimited prompt length (openai_answer_once), risking context window exhaustion.
Split `build_or_load_vectorstore` into two tools: `build_vectorstore(text, file_name)` and `load_vectorstore(file_name)`. Add a `confirm_rebuild` parameter to prevent accidental overwrites. Document that building writes to disk under faiss_indexes/<sanitized_name>/.
Remove or hide utility tools (`get_index_folder`, `ensure_chunks`). These are internal helpers, not agent-level operations. Encapsulate them inside `build_vectorstore`.
Make tool signatures portable. Replace Streamlit-specific inputs (e.g., `pdf_file` as uploader object) with standard formats: base64-encoded bytes, file URI, or inline content. Example: `read_pdf_text(pdf_content: base64 bytes) -> text: string`. Remove `@st.cache_resource` and `st.session_state` references; use standard Python caching or a database.
Add parameter constraints and validation. For `ensure_chunks`, document: 'chunk_size=800, chunk_overlap=150, min_chunk_length=80 chars, filters URLs and links.' For `free_answer`, enforce max context length: 'context maxLength=6000'. For `openai_answer_*`, set prompt max length: 'prompt maxLength=8000'.
Document downstream dependencies. Make clear which tools must be called in sequence. Example: 'RAG workflow: read_pdf_text() → ensure_chunks() → build_vectorstore() → [free_answer() or openai_answer_stream()] with retrieved context.' Update descriptions to guide this order.
Add idempotency and dry-run support. For `build_vectorstore`, accept a `dry_run=true` parameter to validate inputs without writing to disk. Ensure repeated calls with same inputs return identical results (or skip rebuild if vectorstore already exists).
Return comprehensive schemas with chaining IDs. Example: `read_pdf_text()` returns {"text": "...", "page_count": 10, "file_name_sanitized": "paper_xyz"}. Include `file_name_sanitized` so caller can immediately pass it to `build_vectorstore(file_name)` without extra lookup.
Document rate limits and timeouts. Add guidance: 'OpenAI tools are rate-limited to X calls/min. If you hit a rate limit error, retry after 60 seconds or use `free_answer()` instead.' Set explicit timeouts on API calls (e.g., 30s for OpenAI, 10s for local models).
Add tool annotations (readOnlyHint, destructiveHint, idempotentHint) for MCP 2026-07-28 spec alignment. Example: `build_or_load_vectorstore` should have destructiveHint=true (overwrites existing vectorstore). `free_answer` and read-only tools should have readOnlyHint=true.