Data Engineering - Research Assistant
Job Description SummaryOrganization's Summary Statement: The Applied Research Laboratory for Intelligence & Security (ARLIS) at the University of Maryland is a University-Affiliated Research Center (UARC) dedicated to advancing research, innovation, and technology transition to improve decision making for U.S. national security. ARLIS combines deep scientific expertise with operational insight to address challenges in intelligence analysis, cybersecurity, artificial intelligence / machine learning, quantum science, and human-machine teaming. Researchers, scientists, engineers, and analysts at ARLIS collaborate with government agencies, industry partners, and academic institutions to deliver actionable insights and transformative solutions through research and development. Employees at ARLIS work on projects of critical importance, contribute directly to the nation’s security, and are supported by a culture that values integrity, collaboration, and professional growth. The Data Engineering Research Assistant [TW2.1]works directly with project teams to build and run Python-based data pipelines, prepare, load, and validate data in databases, and help implement search and retrieval components under the guidance of senior technical staff.Physical Demands: Sedentary work performed in a normal office environment; exerts up to 10 pounds of force occasionally and/or negligible amount of force frequently or constantly to lift, carry, push, pull or otherwise move objects, including the human body. Ability to attend meetings both on and off campus. Spending long hours in front of a computer screen.Licenses/ Certifications: N/AMinimum QualificationsEducation:Bachelor’s degree in Computer Science, Data Science, Information Systems, Engineering, Mathematics, or a related field from an accredited college or university.Experience:Professional or substantial project-based experience programming in Python.Experience designing and implementing data pipelines or ETL/ELT jobs that read from one or more sources, transform data, and load into a database.Hands-on experience with at least one database (document or relational), including schema design, writing queries, and validating loaded data.Experience using Linux command-line tools and working in a Git-based workflow (branches and pull requests).Experience writing and executing tests in Python (e.g., using Pytest).Experience using Docker or similar containerization tools to run or troubleshoot applications.Knowledge, Skills, and Abilities:Ability to quickly internalize and work within modular codebases and design re-runnable, configurable jobs.Conceptual understanding of embeddings and vector similarity search, or prior experience with information retrieval / NLP projects.Familiarity with tools and technologies such as MongoDB, PostgreSQL, JSON Schema, Qdrant, pgvector, FAISS, or sentence-transformers (experience with some subset is acceptable).Ability to build or confidently learn to build REST APIs with FastA...