Joseph Junior Mensah
VectorNexus
← BACK TO PROJECTS
RAGFull-StackAI + Cloud2026

◎VectorNexus

A Cornell-gated retrieval assistant over Confluence: hybrid dense + sparse search with reranking, every answer cited to its source, Claude on Bedrock via an IAM role (no keys in the box), shipped to AWS with a zero-SSH pipeline.

Timeline

2026

Team

Solo build

Role

Design & Full-Stack Engineering

Skills

RAG, Full-Stack, AI + Cloud

Built with

Next.js·TypeScript·Python·FastAPI·LangChain·Qdrant·fastembed·Claude (Amazon Bedrock)·Docker·AWS (EC2 · ECR · SSM · Bedrock)·Caddy·Resend

LONG  STORY  SHORT

I built a retrieval assistant for Cornell teams that answers from their Confluence and shows its receipts — hybrid search, on-device embeddings, and Claude on Bedrock with no API key anywhere in the box.

Ask a large-language model about your team's internal docs and it will answer confidently whether or not it actually knows. VectorNexus refuses to do that. It answers from a specific knowledge base — a team's Confluence — and hands back the passages it used, numbered and clickable, so a reader can check the work instead of trusting it. It's built for Cornell project teams, so the front door is a Cornell-only gate: a NetID gets you a one-time code, and nothing else gets through.

Every answer hands back the passages it used — numbered, quoted, and clickable. No confident guessing.
Every answer hands back the passages it used — numbered, quoted, and clickable. No confident guessing.
Sources come back numbered; each one links to the exact Confluence page it came from.
Sources come back numbered; each one links to the exact Confluence page it came from.

How it works

A question isn't answered from the model's memory — it's answered from what's actually indexed. Each Confluence page is cleaned to text, split into overlapping chunks, and embedded into a vector store. At query time VectorNexus runs hybrid retrieval (dense semantic search and sparse keyword search in parallel), fuses the two rankings, reranks the survivors with a cross-encoder, and hands only the top passages to the model — which then answers and streams back the citations. Every embedding, fusion, and rerank step runs on the box itself; only the final generation calls out to a model.

The retrieval pipeline

One question, end to end — from a NetID-gated prompt to a streamed answer whose every claim points at a source passage.

  1. 1

    Ask

    gated by Cornell NetID

  2. 2

    Hybrid retrieve

    dense + BM25

  3. 3

    Fuse

    reciprocal rank fusion

  4. 4

    Rerank

    cross-encoder

  5. 5

    Answer

    Claude on Bedrock

  6. 6

    Cite

    numbered sources

One question, end to end — from a NetID-gated prompt to a streamed answer that cites the passages it used.

Architecture

One small box, five clean tiers. The browser only ever talks to Caddy, which terminates TLS and routes /api/* to the FastAPI backend and everything else to the Next.js app — so the backend and vector store are never exposed to the internet. The backend owns retrieval and the agentic answer, calling Claude on Amazon Bedrock through the instance's IAM role.

Client

Next.js

  • Cornell NetID gate (OTP)
  • Streaming chat workspace
  • Served behind Caddy TLS
⇄HTTPS · /api via Caddy

Server

FastAPI + Qdrant

  • Hybrid retrieval + rerank
  • On-device ONNX embeddings
  • HMAC session, gated chat
⇄InvokeModel (IAM role)

Model

Amazon Bedrock

  • Claude Haiku via IAM role
  • No API key on the box
  • Streamed generation
One box, three tiers — the browser streams from a FastAPI server that retrieves locally and calls Claude on Bedrock through an IAM role, all behind Caddy.

Hybrid search, on-device

The résumé version of "vector search" is one embedding model and cosine similarity. That misses exact terms — a part number, an acronym, a function name. VectorNexus runs dense and sparse retrieval and fuses them with reciprocal rank fusion, then reranks with a cross-encoder for precision. All three models run locally via ONNX, so there's no embedding API key and no per-query embedding cost — the only key the system needs is for generation, and on Bedrock even that isn't a key.

bge-small-en

Dense semantic embeddings.

384-dim · ONNX · on-device

BM25

Sparse keyword scoring for exact terms.

acronyms · part numbers · code

RRF fusion

Merges the dense and sparse rankings.

reciprocal rank fusion, in Qdrant

cross-encoder

Reranks the fused candidates for precision.

MiniLM · top passages to the model

The whole retrieval stack runs on the box — no embedding API key, no per-query embedding cost.

The Cornell gate

Access is a stateless, passwordless gate. A NetID maps to a @cornell.edu address — the address field is split so a non-Cornell domain is something you can't type, not just something that's rejected. A six-digit code is emailed (via Resend), hashed and stored with a ten-minute TTL and a resend cooldown; verifying it mints an HMAC-signed session cookie with no server-side session store to look up. The assistant endpoint independently verifies that signature on every call, so the gate protects the data, not just the page.

  1. 1

    NetID

    @cornell.edu, locked

  2. 2

    Email a code

    Resend · 6 digits · 10 min

  3. 3

    Verify

    hashed · single-use

  4. 4

    Signed session

    HMAC cookie, no store

  5. 5

    Enter

    chat re-checks the signature

Passwordless and stateless — a NetID, an emailed code, and a signed cookie the assistant re-verifies on every call.
The landing page's scalloped ink footer — built at Cornell, for Cornell.
The landing page's scalloped ink footer — built at Cornell, for Cornell.

Shipping it

The whole thing runs as four containers on a single EC2 box — Caddy, the Next.js app, FastAPI, and Qdrant — behind an Elastic IP with automatic HTTPS. There is no API key anywhere on the box: Claude is reached through Amazon Bedrock using the instance's IAM role, scoped to a single Haiku model. Deploys are a git push — GitHub Actions builds both images, pushes to ECR, and the box pulls them via AWS SSM. No SSH, port 22 closed, no static credentials — GitHub authenticates with OIDC and the box uses its own instance role. Stopping the instance pauses the bill.

Why I built it this way

  • Cite everything, guess nothing. Retrieval hands the model real passages and the answer streams back with numbered, clickable sources — a reader verifies rather than trusts.
  • Hybrid beats single-vector. Dense semantics and sparse keywords, fused and reranked, so exact terms and paraphrases both land.
  • Free where it can be. Embedding and reranking run on-device via ONNX — no embedding key, no per-query cost.
  • No keys in the box. Claude runs on Bedrock through an IAM role, so the most sensitive credential simply doesn't exist on the server.
  • The gate protects the data, not the page. A signed session is re-verified on the assistant endpoint itself, so a forged cookie loads an inert shell and nothing more.
  • Boring to operate. One box, one git push to deploy over a zero-SSH OIDC → ECR → SSM pipeline, and a stop button that pauses the bill.
The verify step: a six-digit code, ten-minute TTL, single-use — then a signed session.
The verify step: a six-digit code, ten-minute TTL, single-use — then a signed session.

NEXT  UP  …