Young Investigator, Open Language Models for Biology
Persons in these roles are expected to work from our offices in Seattle. On-site requirements vary based on position and team. If you have questions about on-site work arrangements for this role, please ask your recruiter. Compensation: $159,650.00 Who You Are: Ai2 is seeking a Postdoctoral Young Investigator to work on CellOLMo, a collaboration between Ai2’s open language model researchers and scientists at the Allen Institute. We are looking for a researcher who works at the intersection of machine learning and biology. You should have experience training or adapting transformer-based models and an interest in applying those methods to single-cell data. You may come from machine learning, computational biology, neuroscience, physics, or another quantitative field. You will be expected to design careful experiments, work independently, and communicate with researchers from different disciplines. Because CellOLMo is an open-science project, you should also be committed to releasing code, data, and model weights for use by the broader research community. Who We Are: CellOLMo is investigating how open language models can be used with single-cell brain data to answer biological questions. The project will initially focus on the SEA-AD Alzheimer’s disease atlas and will examine data at several levels, from individual cells to brain regions and donors. One of our main research questions is whether grounding a model in biological language and knowledge produces measurable improvements over models trained only on molecular data. We will also study whether those improvements persist when cells are combined into representations of brain regions, patients, pathology, and disease progression. You will work directly with the project PI at Ai2 and collaborate with domain scientists at the Allen Institute for Brain Health. You will have access to Ai2’s open language models, post-training infrastructure, and substantial GPU compute. We plan to release the resulting models, code, and research artifacts openly. Your Next Challenge: Develop a multimodal language model that combines single-cell gene-expression data with text, using an open model such as OLMo. Reproduce and evaluate relevant approaches, including scGPT, Geneformer, CellWhisperer, and C2S-Scale. Adapt and improve these approaches for brain and neurodegeneration data. Design controlled experiments to test whether language grounding improves performance over cell-only and text-free baselines. Develop representations at the brain-region and donor levels that incorporate pathology and disease progression. Evaluate models on held-out donors using measures such as cell-type and marker recovery, agreement with known patterns of regional vulnerability, and expert review of plain-language answers. Work with Allen Institute scientists to interpret results and identify predictions suitable for experimental follow-up. Release data, code, and model weights, and prepare the results for publication. What Y...