Applied AI Engineer · OpenAI
Agents, evaluation & applied ML
WesleyPasfield
I work on deploying AI systems at OpenAI. Most of my writing is about agents, evaluation, and what we learn from using them in practice.
Take a look at my work
Selected projects
OpenAI Cookbook · 2026
Building an agent improvement loop
Using traces and evaluation results to identify failures, then working with Codex to implement and test changes.
OpenAI Cookbook · 2026
Keeping agents useful over longer tasks
How memory and compaction help agents retain useful context as a conversation grows.
ACM CAIS · 2026
Steering Agent Behavior via a Domain Expert-Driven Alignment-to-Optimization Bridge
Using domain expert trace labels to calibrate judges, optimize prompts, and version agent artifacts.
Articles & research
More on SubstackArticles
Build an Agent Improvement Loop with Traces, Evals, and Codex
A cookbook for turning traces, feedback, and evals into an agent improvement loop that Codex can help implement.
Building Reliable Agents with Memory and Compaction
A cookbook showing how memory and compaction support reliable long-running agent workflows with the OpenAI Agents SDK.
Parallel Tool Calling Agents on Databricks Model Serving via Managed MCP
Exploring how to build parallel tool-calling AI agents using Databricks Model Serving with Managed MCP.
Agent Optimization Pipeline
A practical guide to building agent optimization loops with domain-specific judges, expert calibration, and automated prompt improvement.
How Databricks Helps Baseball Teams Gain an Edge with Data & AI
Turning pitch data into dugout decisions with Unity Catalog, Agent Bricks, and Lakebase.
Self-Optimizing Football Chatbot Guided by Domain Experts on Databricks
Building a self-optimizing chatbot for football analytics guided by domain expertise.
Pilot to Production: Custom Judges
A guide to developing and deploying custom judges for AI evaluation systems.
AI Agents In Advertising for Contextual Content Placement
An agent that uses multimodal models, Unity Catalog, and vector search to identify contextual advertising placements.
Incorporate Agent Evaluations into Your LLMOps with GitHub Actions
How to integrate agent evaluations into your LLMOps workflow using GitHub Actions.
Streamlining AI Paper Discovery: Building an Automated Research Assistant
A guide to building an automated system for discovering and organizing AI research papers.
Publications
Steering Agent Behavior via a Domain Expert-Driven Alignment-to-Optimization Bridge
Accepted ACM CAIS 2026 demo on using domain expert trace labels to calibrate judges, optimize prompts, and version agent artifacts.
Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences
This paper advocates for using curated data as the primary means of LLM evaluation and regulation, and proposes a certification process to ensure that LLMs are safe for public use. Presented at the 2nd Workshop on Regulatable ML at NeurIPS 2024. View Conference Page
About me
I've spent more than 15 years working in machine learning, data science, and analytics.
I'm now an Applied AI Engineer at OpenAI. My work focuses on deploying AI systems and improving them through evaluation and feedback from domain experts.
Previously, I worked at Databricks and served as an Emerging Technology Fellow at the US Census Bureau, where I focused on AI policy and LLM evaluation. Earlier roles took me through Lark Health, Amazon, AWS, Twitch, GoPro, and Nielsen.
I also teach in the University of San Diego's Applied Artificial Intelligence program.
View my résuméAreas I work in
Agent development & evaluation · Applied machine learning · AI policy · Data science leadership
Career history
Applied AI Engineer
OpenAI
2026 - Present
Specialist Solution Architect: Generative AI and Machine Learning
Databricks
2025 - 2026
Emerging Technology Fellow
US Census Bureau
2024 - 2025
Adjunct Professor (Part-time)
Artificial Intelligence and Data Science, University of San Diego
2022 - Present
Head of Data Science
Lark Health
2022 - 2024
Sr. Analytics Manager and Sr. Data Scientist
Amazon Advertising
2019 - 2022
AWS Professional Services ML Specialty Practice Sr. Data Scientist
Amazon Web Services
2017 - 2019
Data Scientist/Sr. Data Scientist
AWS-owned Twitch
2016 - 2017
Manager of User Insights & Acting Data Product Manager
GoPro
2015 - 2016
Analyst, Sr. Analyst and Manager of Custom Analytics
The Nielsen Company
2011 - 2015