Development is documented in my CV. Detailed operational patterns are design notes; public deployment and performance benchmarks are unverified.
Overview
PolicyLens is an enterprise document assistant developed with source citations, role-based access, tool calling, evaluation pipelines and human approval for sensitive actions.
The notes below describe how those capabilities fit together. Specific security guarantees, evaluation results and production deployment behavior have not been verified.
Problem
An apparently useful answer can still cite an outdated policy, expose a restricted document or infer a rule that the evidence does not support. Enterprise retrieval must preserve source versions and enforce access boundaries before information reaches the model.
System Architecture
- 01DocumentsVersioned sources
- 02IngestionPython
- 03RetrievalVector search · RBAC
- 04Grounded answerLLM · RAG
- 05Evidence checksEvaluation
- 06Application APIFastAPI
- 07Human reviewApproval workflow
A document pipeline preserves source versions, section boundaries and permission metadata while creating a searchable index. The request path authenticates the user, retrieves authorized passages and asks the model for a grounded answer with citations.
Tool requests use a separate action boundary. Consequential operations produce a concrete proposal for human approval; a generated answer alone does not authorize execution.
Engineering Decisions
The design treats the model as an untrusted reasoning component within a controlled application.
- Filter retrieval by authorization before constructing the model context.
- Keep stable citation anchors linked to the exact document version used.
- Support an explicit insufficient-evidence response instead of requiring an answer.
- Recheck authorization when serving citations and before executing an approved action.
- Version prompts, retrieval settings and evaluation datasets together for meaningful comparisons.
Technology
Python handles parsing, indexing and evaluation workflows. Vector search supports semantic retrieval, FastAPI exposes the application boundary, and an LLM generates evidence-grounded responses. Provider and index choices remain configurable design decisions.
Data Flow
An approved document is parsed, chunked and indexed with its source and access metadata. An authenticated query retrieves permitted passages, produces an answer and returns source citations. A proposed tool action follows an auditable approval path before any side effect.
Challenges
Permission changes must reach the index without leaving a stale access window. Chunk boundaries must preserve enough context to avoid misleading citations. Evaluation needs examples with missing, contradictory and superseded evidence, alongside ordinary answerable questions.
Trade-offs
Stricter evidence requirements may cause more abstentions, but that is preferable when the source material cannot support an answer. Larger retrieval contexts can improve coverage while increasing latency and distracting the model. Human approval improves control at the cost of an additional operational step.
Observability
Proposed traces capture retrieval latency, document-version references, model and prompt versions, abstentions and approval outcomes. Evaluation should distinguish retrieval recall, citation support and authorization failures. Sensitive prompts and document content should not be logged by default.
Security
Document text must be treated as data rather than trusted instructions. The design calls for defense against prompt injection, least-privilege tool scopes, server-side authorization and explicit approval of consequential actions. Permission-aware caching and access revalidation are essential to prevent cross-user leakage.
Deployment
The proposed deployment separates the API from indexing work, with model and search credentials provided at runtime. New retrieval or prompt configurations should pass a versioned evaluation set before promotion. No public service availability is claimed.
Results
The documented project combines cited answers, role-based access, tool calling, evaluation and an approval path for sensitive actions. No accuracy scores or live usage metrics are published.
Validation should include citation correctness, retrieval performance, refusal behavior, permission changes and attempted cross-role disclosure. Results need an explicit dataset and evaluation method before being reported.
What I Learned
The central engineering takeaway is that trustworthy RAG is a retrieval and access-control problem as well as a generation problem. The model is one component in a system whose sources, permissions and actions need explicit boundaries. Detailed operational lessons will be added with supporting evaluation evidence.
