Talk: RAG on Rails
Video and notes from a talk I gave at London Ruby User Group on building retrieval pipelines in Rails.
I recently gave a talk at the London Ruby User Group on building a Retrieval Augmented Generation (RAG) pipeline in Ruby on Rails. The talk covered what RAG is, how to build a retrieval pipeline using Postgres and familiar Ruby tooling, and how to structure prompts for high-quality, grounded responses with citations.
Here are the gems I referenced during the talk.
Gems
neighbor
Nearest neighbour search for Rails. Sits on top of pgvector and makes it simple to add vector embeddings as a column on any ActiveRecord model. Supports cosine, Euclidean, and inner product distance functions. If you’re doing RAG in Rails, you’ll probably end up using this extensively.
pg_search
Full-text search for ActiveRecord using PostgreSQL’s built-in search capabilities. Useful alongside vector search for keyword-based retrieval where exact term matching matters — for example, domain-specific terminology that embedding models may not handle well.
pragmatic_segmenter
Rule-based sentence boundary detection that works across many languages. Handles the tricky edge cases of sentence splitting — abbreviations, bullet points, decimal numbers — that naive full-stop splitting gets wrong. Essential for chunking documents at sensible boundaries.
engtagger
English part-of-speech tagger. A probability-based, corpus-trained tagger that assigns POS tags to text and can extract nouns and noun phrases. Useful for pulling keywords out of a search query to feed into keyword-based retrieval alongside semantic search.
Bonus
ruby_llm
One unified Ruby interface for OpenAI, Anthropic, Gemini, Bedrock, Ollama, and many more providers. Includes Rails integration with generators, streaming support, tool use, vision, audio, and embeddings. Three dependencies. If you’re building LLM-powered features in Rails, this is a great starting point.