AI Data Engineer
Job DescriptionAgilent inspires and supports discoveries that advance the quality of life by providing life science, diagnostic and applied market laboratories worldwide with instruments, services, consumables, application and both measurement and asset management expertise.Builds the data products and pipelines a pod runs on; turns raw domain data into model-ready assets. Pods do not wait for the Fabric to be complete; they build the Fabric through execution, and this role is where that happens for the data plane. Every domain data product built in a pod is constructed to certification standards from the start, because the second consumer of the asset is the point, not an afterthought. The role goes beyond the conventional pipeline engineering. AI consumption changes what "model-ready" means: data must be retrieval-ready, semantically annotated, contract-governed, and quality-scored, and the AI Data Engineer often uses agents to do the building, generating metadata, resolving entities, and classifying unstructured domain content rather than hand-curating at a scale that cannot hold. Responsible for Domain data products and pipelines serving the pod's use case, built to the data plane's certification standards: semantic definition, data contract, entitlement metadata including agent identity, lineage, and certification tier from day one. Data quality and model-readiness, including quality scoring against the defined dimensions and the quality signals that feed the evals spine; when quality falls below threshold, this role is the one who says so before the agent does something embarrassing with it. Working relationships with the domain's data owners and stewards, so that steward-validated definitions and certified sources ground the pod's retrieval rather than whatever was easiest to reach. Retrieval foundations for the pod: structured and unstructured grounding, vector and graph assets where the use case requires relationship reasoning, built on the platform estate rather than parallel infrastructure. AI-built curation in practice: deploying metadata generation, entity resolution, and content classification agents against the domain's data, contributing those outputs to the registry. Reusable assets back to the Fabric: every domain data product is a candidate for certification and enterprise reuse, documented and handed to the data plane for promotion, not a point integration that dies with the pod. Works with the business teams and IT to develop domain data models.Applying broad understanding of on premise and cloud deployment topologies, creates data collection frameworks to capture, manage, store and utilize large sets of structured, unstructured and/or disconnected data from a wide variety of internal and external sources.Oversees the establishment of data set processes and builds data structures based on business and technical requirements to funnel data from ...