Thomsonreuters

Lead Applied Scientist, Document Understanding

Full timeLeadPosted about 14 hours ago
Apply on Thomsonreuters →

Sign in to see who you know at Thomsonreuters.

Lead Applied Scientist, Document Understanding

This position is based in either Zug, Switzerland or London, UK.

Want to use your experience of building document understanding AI solutions to enhance our leading products in the tax, legal and professional services industries?

Document understanding is a foundational intelligence layer that powers every major capability across our legal AI platform—from search and information extraction to agentic reasoning in products like Westlaw, PracticalLaw, and CoCounsel. You'll build state-of-the-art semantic chunking, document enrichment, and knowledge graph construction systems that serve as the cognitive foundation multiple product teams depend on, working across authoritative legal, tax, and accounting content and extraordinarily diverse customer data.

This is a rare opportunity to solve publishing-quality research problems with immediate production impact—your innovations will directly shape how millions of legal professionals research, analyze, and reason over complex legal documents while advancing the capabilities that enable the next generation of intelligent legal AI agents.

You will work across semantic chunking, document enrichment, intelligent classification, information extraction, knowledge graph construction, and tabular data understanding for complex legal, tax, and accounting content. This work forms the foundation for downstream search, retrieval, reasoning, and generative AI experiences.

About the Role

As Lead Applied Scientist, Document Understanding at Thomson Reuters, you will:

Design and deploy semantic chunking systems for lengthy, non-uniformly structured legal, tax, and accounting documents

Build document enrichment pipelines that identify document types, jurisdictions, legal concepts, entities, parties, and other domain-specific metadata

Develop hierarchical and multi-label document classification systems using both standard and customer-defined taxonomies

Build LLM-based and traditional NLP information extraction pipelines that identify entities, relationships, citations, references, and key concepts from unstructured content

Develop knowledge graph construction systems that extract, normalize, connect, and enrich entities, legal concepts, citations, and relationships across large document collections

Design systems that identify, interpret, and extract insights from complex tabular data embedded within legal, tax, regulatory, and accounting documents

Create document intelligence capabilities that support downstream search, retrieval, RAG, and agentic AI workflows

Design robust evaluation frameworks for document understanding systems using expert annotations, synthetic datasets, and production metrics

Lead technical decisions on document analysis architectures, chunking strategies, extraction methodologies, classification approaches, and knowledge representation frameworks

Partner closely with engineering teams to deliver...