
◎VectorNexus
A Cornell-gated retrieval assistant over Confluence: hybrid dense + sparse search with reranking, every answer cited to its source, Claude on Bedrock via an IAM role (no keys in the box), shipped to AWS with a zero-SSH pipeline.
Timeline
2026
Team
Solo build
Role
Design & Full-Stack Engineering
Skills
RAG, Full-Stack, AI + Cloud
Built with
LONG STORY SHORT
I built a retrieval assistant for Cornell teams that answers from their Confluence and shows its receipts — hybrid search, on-device embeddings, and Claude on Bedrock with no API key anywhere in the box.
Ask a large-language model about your team's internal docs and it will answer confidently whether or not it actually knows. VectorNexus refuses to do that. It answers from a specific knowledge base — a team's Confluence — and hands back the passages it used, numbered and clickable, so a reader can check the work instead of trusting it. It's built for Cornell project teams, so the front door is a Cornell-only gate: a NetID gets you a one-time code, and nothing else gets through.


How it works
A question isn't answered from the model's memory — it's answered from what's actually indexed. Each Confluence page is cleaned to text, split into overlapping chunks, and embedded into a vector store. At query time VectorNexus runs hybrid retrieval (dense semantic search and sparse keyword search in parallel), fuses the two rankings, reranks the survivors with a cross-encoder, and hands only the top passages to the model — which then answers and streams back the citations. Every embedding, fusion, and rerank step runs on the box itself; only the final generation calls out to a model.
The retrieval pipeline
One question, end to end — from a NetID-gated prompt to a streamed answer whose every claim points at a source passage.
- 1
Ask
gated by Cornell NetID
- 2
Hybrid retrieve
dense + BM25
- 3
Fuse
reciprocal rank fusion
- 4
Rerank
cross-encoder
- 5
Answer
Claude on Bedrock
- 6
Cite
numbered sources
Architecture
One small box, five clean tiers. The browser only ever talks to Caddy, which terminates TLS and routes /api/* to the FastAPI backend and everything else to the Next.js app — so the backend and vector store are never exposed to the internet. The backend owns retrieval and the agentic answer, calling Claude on Amazon Bedrock through the instance's IAM role.
Client
Next.js
- Cornell NetID gate (OTP)
- Streaming chat workspace
- Served behind Caddy TLS
Server
FastAPI + Qdrant
- Hybrid retrieval + rerank
- On-device ONNX embeddings
- HMAC session, gated chat
Model
Amazon Bedrock
- Claude Haiku via IAM role
- No API key on the box
- Streamed generation
Hybrid search, on-device
The résumé version of "vector search" is one embedding model and cosine similarity. That misses exact terms — a part number, an acronym, a function name. VectorNexus runs dense and sparse retrieval and fuses them with reciprocal rank fusion, then reranks with a cross-encoder for precision. All three models run locally via ONNX, so there's no embedding API key and no per-query embedding cost — the only key the system needs is for generation, and on Bedrock even that isn't a key.
bge-small-en
Dense semantic embeddings.
384-dim · ONNX · on-device
BM25
Sparse keyword scoring for exact terms.
acronyms · part numbers · code
RRF fusion
Merges the dense and sparse rankings.
reciprocal rank fusion, in Qdrant
cross-encoder
Reranks the fused candidates for precision.
MiniLM · top passages to the model
The Cornell gate
Access is a stateless, passwordless gate. A NetID maps to a @cornell.edu address — the address field is split so a non-Cornell domain is something you can't type, not just something that's rejected. A six-digit code is emailed (via Resend), hashed and stored with a ten-minute TTL and a resend cooldown; verifying it mints an HMAC-signed session cookie with no server-side session store to look up. The assistant endpoint independently verifies that signature on every call, so the gate protects the data, not just the page.
- 1
NetID
@cornell.edu, locked
- 2
Email a code
Resend · 6 digits · 10 min
- 3
Verify
hashed · single-use
- 4
Signed session
HMAC cookie, no store
- 5
Enter
chat re-checks the signature

Shipping it
The whole thing runs as four containers on a single EC2 box — Caddy, the Next.js app, FastAPI, and Qdrant — behind an Elastic IP with automatic HTTPS. There is no API key anywhere on the box: Claude is reached through Amazon Bedrock using the instance's IAM role, scoped to a single Haiku model. Deploys are a git push — GitHub Actions builds both images, pushes to ECR, and the box pulls them via AWS SSM. No SSH, port 22 closed, no static credentials — GitHub authenticates with OIDC and the box uses its own instance role. Stopping the instance pauses the bill.
Why I built it this way
- Cite everything, guess nothing. Retrieval hands the model real passages and the answer streams back with numbered, clickable sources — a reader verifies rather than trusts.
- Hybrid beats single-vector. Dense semantics and sparse keywords, fused and reranked, so exact terms and paraphrases both land.
- Free where it can be. Embedding and reranking run on-device via ONNX — no embedding key, no per-query cost.
- No keys in the box. Claude runs on Bedrock through an IAM role, so the most sensitive credential simply doesn't exist on the server.
- The gate protects the data, not the page. A signed session is re-verified on the assistant endpoint itself, so a forged cookie loads an inert shell and nothing more.
- Boring to operate. One box, one
git pushto deploy over a zero-SSH OIDC → ECR → SSM pipeline, and a stop button that pauses the bill.

NEXT UP …


