Skip to Main Content
Talk Intermediate BSD-3-Clause license First Talk

Seqout: search engine for the world’s genomic data

Proposal status is Approved
Session Description

DNA is the molecular blueprint of life. Modern sequencing technologies allow us to read DNA and RNA at unprecedented scale, generating vast amounts of biological data from organisms, tissues, and individual cells. Today, public repositories host millions of sequencing experiments comprising petabytes of data that power discoveries in medicine, drug development, agriculture, and fundamental biology.

Yet finding the right dataset is often harder than analysing it. Researchers must navigate fragmented repositories, heterogeneous metadata standards, and incomplete descriptions, making dataset discovery a time-consuming and frustrating process. With Seqout (https://seqout.org), we set out to build a genomics search engine that is fast, comprehensive, and lean enough to run on modest hardware while serving hundreds of thousands of researchers worldwide. Seqout is built entirely on open-source software: PostgreSQL and Neo4j power keyword and graph-based search, while embedding-based retrieval is enabled by open-source models and libraries from Hugging Face.

In this talk, we will discuss the engineering challenges involved in building a low-latency search engine over millions of biological datasets with limited resources (aka compute and GPUs). We will cover data ingestion, metadata normalisation, indexing strategies, typo tolerance, semantic search, and the trade-offs involved in integrating relational and graph databases with vector search.

Along the way, we'll share lessons learned from operating search infrastructure at scale and discuss how open-source software powers every layer of the platform. We'll also discuss how Seqout's MCP server enables large language models and AI agents to move beyond answering biology questions to directly discover and interact with relevant scientific datasets.

Key Takeaways
  1. How to build a low-latency search engine from scratch
  2. Combining PostgreSQL, Neo4j, and semantic search to improve dataset discovery
  3. Lessons learned from building and operating an open-source scientific search platform

References

Session Categories

Introducing a FOSS project or a new version of a popular project
Talk License: BSD-3-Clause license

Which track are you applying for?

Main track

Speakers

Aniruddha Mukherjee Graduate student | Indian Institute of Technology Bombay

I am a master's student at IIT Bombay, developing tools and methods for computational biology at Saket Lab at the Koita Centre for Digital Health.

Aniruddha Mukherjee
https://amkhrjee.in/
Mukesh Reddy Undergraduate Student | NIT Warangal

I'm an undergraduate student pursuing a Bachelor's degree in Biotechnology at NIT Warangal. I worked as an undergraduate research intern at the Koita Centre for Digital Health at IIT Bombay, where I contributed to the development of Seqout.

Mukesh Reddy
https://0xmukesh.github.io