Key Responsibilities
1. AI Agent Architectures
- Design and develop modern AI agent architectures.
- Work with planning, tool calling, function calling, memory, and context engineering.
- Implement semantic routing and determine when to use workflows vs. autonomous agents.
- Work with Model Context Protocol (MCP) and, ideally, Agent-to-Agent (A2A) architectures.
- Build and maintain multi-agent orchestration systems.
- Work with frameworks such as LangGraph and OpenAI Agents SDK.
2. Agent Runtime Engineering
- Design and implement robust agent execution loops and runtimes.
- Develop state management and orchestration mechanisms.
- Implement retries, checkpointing, streaming, and human-in-the-loop workflows.
- Build reliable tool-execution pipelines.
- Handle context management and concurrency.
- Implement comprehensive observability and tracing for agent execution.
- Improve runtime reliability, scalability, latency, and fault tolerance.
3. AI Evaluation & Evals — Critical Responsibility
- Design and implement offline and online evaluation frameworks for AI systems.
- Develop meaningful benchmarks and golden datasets.
- Generate and manage synthetic evaluation datasets.
- Implement LLM-as-a-Judge and pairwise comparison methodologies.
- Build automated regression-testing suites for LLM applications.
- Establish experiment tracking and ensure reproducibility.
- Analyze model and agent performance to continuously improve quality.
- Work with MLflow for experiment tracking and evaluation workflows.
4. RAG Engineering
- Design and optimize Retrieval-Augmented Generation (RAG) pipelines.
- Work on chunking strategies, embeddings, and reranking.
- Work with vector databases and retrieval systems.
- Measure and improve retrieval quality.
- Optimize context selection and context-window utilization.
5. LLM Engineering
- Develop effective prompting strategies.
- Implement reliable structured outputs and JSON Schema.
- Build and integrate function-calling workflows.
- Understand tokenization, context windows, and reasoning models.
- Optimize LLM applications for cost and latency.
- Implement caching strategies to improve performance and reduce inference costs.
6. Backend & High-Performance Engineering
- Develop scalable and maintainable backend services using Python.
- Design high-performance APIs and distributed backend systems.
- Build production-grade services using technologies such as:
- FastAPI
- Docker
- Kubernetes
- Redis
- Kafka
- PostgreSQL
- Optimize backend systems for throughput, concurrency, reliability, and low latency.
7. Observability & Monitoring
- Implement comprehensive observability across AI and backend systems.
- Work with:
- OpenTelemetry
- Distributed tracing
- Metrics
- Logs
- Agent execution traces
- Cost monitoring
- Latency monitoring
- Experience with Grafana and Prometheus is highly desirable.
8. Testing & Quality Engineering
- Develop comprehensive unit and integration tests.
- Implement AI-specific evaluation testing.
- Build regression test suites for LLM and agent applications.
- Establish automated testing and quality gates.
- Implement CI/CD pipelines for AI/LLM applications.
- Ensure reproducibility and reliability of AI experiments and production systems.
9. Cloud & AI Platforms
Experience with at least one major cloud platform is desirable:
- Microsoft Azure
- AWS
- Google Cloud Platform (GCP)
Particularly valuable experience includes:
- Azure OpenAI
- Google Vertex AI
- Amazon Bedrock
- Amazon Bedrock AgentCore
Required Skills & Experience
- 4+ years of professional software development experience.
- Strong and demonstrable expertise in Python — mandatory.
- Strong backend engineering fundamentals.
- Experience building production-grade, scalable software systems.
- Solid understanding of AI/LLM application architecture.
- Hands-on experience with agent systems, LLM evaluation, or RAG.
- Strong understanding of APIs, distributed systems, concurrency, and asynchronous processing.
- Experience with testing, debugging, monitoring, and production troubleshooting.
- Ability to work independently on technically complex engineering problems.
Highly Desirable Skills
- LangGraph
- OpenAI Agents SDK
- MCP (Model Context Protocol)
- A2A
- MLflow
- FastAPI
- Docker & Kubernetes
- Redis
- Kafka
- PostgreSQL
- OpenTelemetry
- Grafana & Prometheus
- Azure OpenAI
- Vertex AI
- Amazon Bedrock / AgentCore