DNA is the molecular blueprint of life. Modern sequencing technologies allow us to read DNA and RNA at unprecedented scale, generating vast amounts of biological data from organisms, tissues, and individual cells. Today, public repositories host millions of sequencing experiments comprising petabytes of data that power discoveries in medicine, drug development, agriculture, and fundamental biology.
Yet finding the right dataset is often harder than analysing it. Researchers must navigate fragmented repositories, heterogeneous metadata standards, and incomplete descriptions, making dataset discovery a time-consuming and frustrating process. With Seqout (https://seqout.org), we set out to build a genomics search engine that is fast, comprehensive, and lean enough to run on modest hardware while serving hundreds of thousands of researchers worldwide. Seqout is built entirely on open-source software: PostgreSQL and Neo4j power keyword and graph-based search, while embedding-based retrieval is enabled by open-source models and libraries from Hugging Face.
In this talk, we will discuss the engineering challenges involved in building a low-latency search engine over millions of biological datasets with limited resources (aka compute and GPUs). We will cover data ingestion, metadata normalisation, indexing strategies, typo tolerance, semantic search, and the trade-offs involved in integrating relational and graph databases with vector search.
Along the way, we'll share lessons learned from operating search infrastructure at scale and discuss how open-source software powers every layer of the platform. We'll also discuss how Seqout's MCP server enables large language models and AI agents to move beyond answering biology questions to directly discover and interact with relevant scientific datasets.