Turn scattered operational documents into grounded, conversational answers.
Enterprise Operations
AWS BedrockClaudepgvectorTitan Embeddings
Knowledge AssistantRAG
Find the maintenance procedure.↵
✧
Source documents
RETRIEVE → GROUND → RESPOND
MY ROLE
AI solution design · RAG implementation · conversational UX
PROJECT STATUS
Working prototype
FOCUS
RAG · Enterprise AI
01 / BUSINESS CHALLENGE
Start with the problem.
Operational knowledge lives across PDFs, manuals, and engineering procedures. Finding the right passage requires searching multiple documents and judging which source is relevant.
02 / SOLUTION APPROACH
A practical path forward.
A retrieval-augmented assistant brings document ingestion, semantic search, and cited responses into one conversational experience.
My contribution
AI solution design · RAG implementation · conversational UX.
03 / ARCHITECTURE
How the pieces connect.
Documented project pattern
Knowledge into grounded answers
1 / 04
Enterprise Knowledge AI AssistantClick a component
Indexing and answering are separate pipelines. Source metadata travels with retrieved passages into the response.
04 / IMPLEMENTATION
From design to workflow.
01
Index the knowledge
Source documents in S3 are read by a Node.js ingestion pipeline and split into approximately 500–700 token chunks. Amazon Titan Embeddings V2 creates vectors, which are stored alongside text and metadata in RDS PostgreSQL with pgvector.
02
Retrieve useful context
At query time, the question is embedded with the same model. pgvector retrieves the five most relevant sections using an HNSW index. Source metadata stays attached to the retrieved context.
03
Generate with evidence
Claude Sonnet, accessed through AWS Bedrock, receives the retrieved passages and response instructions. The answer includes source references so users can inspect its evidence. Retrieval improves grounding; it does not guarantee correctness.
Design decisions & tradeoffs
Why PostgreSQL + pgvector?+
It keeps chunks, source metadata, and embeddings in one familiar database. A dedicated vector service may become useful at a different scale or with different operational requirements.
Why HNSW?+
Approximate nearest-neighbor retrieval without an index training step. Memory use, recall, and index tuning remain tradeoffs.
Why 500–700 token chunks?+
The documented prototype uses this range to balance context with retrieval precision. It is a starting point to evaluate, not a universal optimum.
08 / OUTCOMES
What the work demonstrates.
Implemented conversational retrieval over operational documents.
Connected generated responses to source citations.
Created a reusable approach to ingestion and semantic retrieval.
Working prototype documented in the original portfolio. The public demo uses fictional documents and preset responses. No production deployment or measured business impact is claimed.
09 / INTERACTIVE DEMONSTRATION
Experience the pattern.
SIMULATION
01THE REQUEST
Ask. Retrieve. Verify.
Choose a sample question to explore a grounded response.
Preset responses for the sample topics shown above.
Fictional data. No live AI calls or enterprise connections.
02INSIDE THE WORKFLOW
01Understand
02Retrieve
03Validate
04Generate
05Cite sources
See the workflow unfold.
Run a sample to explore each step, the structured data, and the resulting action.
AWAITING SAMPLE INPUT
10 / REFLECTIONS
What I take forward.
Chunk boundaries must preserve enough context for useful retrieval.
A citation is an inspection path, not proof that an answer is correct.
Keeping vectors and metadata together simplifies the prototype; scaling and access controls need separate evaluation.