Sr Manager Software Dev, Advertising Full Funnel Agentic Intelligence
Job description
Senior Software Development Manager, FAIM Evaluations.
Amazon Advertising is building toward a future where an advertiser specifies a few marketing parameters (budget, success definition, which products to promote) and a set of AI agents handles the rest. The Full Funnel Agentic Intelligence and Models (FAIM) organization owns that bet: the Ads Nova agent, the Ads Nova model it reasons with, and the agent infrastructure, learning environments, and evaluations that connect the two. We are looking for a Senior Software Development Manager to found and lead the FAIM Evaluations team. You will report directly to the Vice President of Full Funnel Agentic Intelligence and Models and own how the entire organization answers one question: is this actually good at advertising?
This is a ground-up build. Evaluation today lives inside individual model and agent teams, measured task by task. You will create the standalone engineering team that turns it into a shared, rigorous system spanning the Ads Nova model, the Ads Nova agent, and the internal agent that serves our Sales, Services, and Operations teams. No inherited harness, no pattern to follow, and a direct line to the VP who sponsors the work.
What we're building
An advertising benchmark: a representative set of real advertising tasks, organized by domain and difficulty, from single-step questions through multi-step analysis to long-horizon strategic work, each with structured criteria for what a correct end-to-end response looks like
Evaluation infrastructure that scores models and agents deterministically against that benchmark, compares Ads Nova to frontier models on the tasks that matter to advertisers, and gives every science and product team in FAIM the same yardstick
Rubrics and task environments built to serve double duty: scoring quality today and producing the reward signal that trains the next version of the model
An expert-in-the-loop program that captures how experienced advertising practitioners actually work and encodes that judgment into criteria a machine can grade against
A capability map, derived from benchmark results, that tells FAIM where the model and agent stand and what to train next
Key job responsibilities
Found and lead a standalone team of roughly 10-12: software engineers plus a product manager and a technical program manager; you will hire most of them
Own the technical vision and roadmap for FAIM evaluations end to end: task taxonomy, rubric design, environment construction, scoring, benchmark versioning, and the separation between what we evaluate on and what we train on
Build for the whole org, not one product: your team evaluates the Ads Nova model, the Ads Nova agent, and the internal agent, and you participate in the planning and reviews for all three
Partner with applied scientists across FAIM to turn evaluation criteria into training signal for reinforcement learning, and to make sure what we measure is what we optimize
Run the domain-expert program: source advertising practitioners, define the annotation and calibration process, and hold the quality bar on inter-rater agreement
Set the evaluation standard for the organization and hold the line on it; where good internal assets already exist, adopt them rather than rebuild
Publish results leadership and partner teams trust, and own the cadence for re-scoring as models, agents, and tasks evolve
Represent evaluations in VP-level reviews, annual planning, and cross-org discussions on model and agent quality
We're looking for a leader who brings
10+ years of engineering experience and 5+ years managing engineering teams, including building a team from a small core
A track record delivering evaluation systems, benchmarks, or data-quality programs for machine learning models, ideally large language models or agentic systems
Working fluency in how modern models are trained and improved (supervised fine-tuning, reinforcement learning from rubric or verifier signal) and what makes an eval useful as a training asset rather than only a scorecard
Judgment about measurement: when all-or-nothing grading beats partial credit, how to find ambiguous criteria through grader disagreement
Experience running expert-annotation or labeling programs with external partners, including quality control at scale
The ability to operate in ambiguity: turn "is it good at advertising" into a concrete, scored, versioned asset with minimal scoping help
Comfort working as the engineering counterpart to scientists you do not manage, and the influence to get model, agent, and product teams onto one yardstick
Advertising domain knowledge, or the curiosity and speed to build it by working closely with practitioners
Experience with Amazon Bedrock, agent frameworks, and tool-use protocols (MCP) is a plus
Skills mentioned
Amazon is a global leader in e-commerce, retail, and cloud computing, dedicated to being the most customer-centric company in the world. The company offers a vast online marketplace, innovative digital content, and cutting-edge cloud services through Amazon Web Services (AWS), serving millions of consumers and businesses worldwide. Amazon fosters a dynamic, fast-paced environment focused on invention and operational excellence. Professionals joining Amazon contribute to groundbreaking technologies like Alexa, while upholding a commitment to long-term thinking and a unique 'Day 1' mentality – prioritizing agility, innovation, and customer delight.
Apply for this job
Use the application link supplied with this listing to apply to Amazon. Check the destination before entering personal information.
