Retrieval over the job descriptions I actually applied to.
Hybrid retrieval. Dense embeddings match "owns the system end to
end" against "full ownership of the stack" where no words overlap. They are
bad at rare exact tokens, because LangGraph,
pgvector and MCP sit near each other and near a
hundred other tool names. Postgres full-text is the reverse. Both run, and
the rankings fuse with Reciprocal Rank Fusion, which needs only each
retriever's ordering and so avoids weighting cosine distance against
ts_rank.
It grades its own retrieval. Retrieve-then-generate fails quietly: if retrieval misses, the model still writes a confident paragraph out of training data. Here a grader reads the chunks first and decides whether they can answer the question at all.
retrieve -> grade -+-- relevant ------> generate -> END
+-- not relevant --> rewrite -> retrieve (max 2)
+-- out of attempts -> admit_gap -> END
The admit_gap branch is the point. Ask it something the
corpus does not cover and it says so instead of inventing an answer.
FastAPI and LangGraph on Vercel, Postgres with pgvector on Supabase.
Embeddings run on gte-small inside a Supabase Edge Function,
because Groq has no embeddings endpoint and retrieval needs one anyway.
Llama 3.3 70B on Groq generates and grades. API health