Detalles Equipo Calendario Documento FAQ
Challenge

Senior Python Developer – AI Runtime & Evaluation Engineer

★
★
★
★
★

Ranking: 2637

Key Responsibilities

1. AI Agent Architectures

  • Design and develop modern AI agent architectures.
  • Work with planning, tool calling, function calling, memory, and context engineering.
  • Implement semantic routing and determine when to use workflows vs. autonomous agents.
  • Work with Model Context Protocol (MCP) and, ideally, Agent-to-Agent (A2A) architectures.
  • Build and maintain multi-agent orchestration systems.
  • Work with frameworks such as LangGraph and OpenAI Agents SDK.

2. Agent Runtime Engineering

  • Design and implement robust agent execution loops and runtimes.
  • Develop state management and orchestration mechanisms.
  • Implement retries, checkpointing, streaming, and human-in-the-loop workflows.
  • Build reliable tool-execution pipelines.
  • Handle context management and concurrency.
  • Implement comprehensive observability and tracing for agent execution.
  • Improve runtime reliability, scalability, latency, and fault tolerance.

3. AI Evaluation & Evals — Critical Responsibility

  • Design and implement offline and online evaluation frameworks for AI systems.
  • Develop meaningful benchmarks and golden datasets.
  • Generate and manage synthetic evaluation datasets.
  • Implement LLM-as-a-Judge and pairwise comparison methodologies.
  • Build automated regression-testing suites for LLM applications.
  • Establish experiment tracking and ensure reproducibility.
  • Analyze model and agent performance to continuously improve quality.
  • Work with MLflow for experiment tracking and evaluation workflows.

4. RAG Engineering

  • Design and optimize Retrieval-Augmented Generation (RAG) pipelines.
  • Work on chunking strategies, embeddings, and reranking.
  • Work with vector databases and retrieval systems.
  • Measure and improve retrieval quality.
  • Optimize context selection and context-window utilization.

5. LLM Engineering

  • Develop effective prompting strategies.
  • Implement reliable structured outputs and JSON Schema.
  • Build and integrate function-calling workflows.
  • Understand tokenization, context windows, and reasoning models.
  • Optimize LLM applications for cost and latency.
  • Implement caching strategies to improve performance and reduce inference costs.

6. Backend & High-Performance Engineering

  • Develop scalable and maintainable backend services using Python.
  • Design high-performance APIs and distributed backend systems.
  • Build production-grade services using technologies such as:
    • FastAPI
    • Docker
    • Kubernetes
    • Redis
    • Kafka
    • PostgreSQL
  • Optimize backend systems for throughput, concurrency, reliability, and low latency.

7. Observability & Monitoring

  • Implement comprehensive observability across AI and backend systems.
  • Work with:
    • OpenTelemetry
    • Distributed tracing
    • Metrics
    • Logs
    • Agent execution traces
    • Cost monitoring
    • Latency monitoring
  • Experience with Grafana and Prometheus is highly desirable.

8. Testing & Quality Engineering

  • Develop comprehensive unit and integration tests.
  • Implement AI-specific evaluation testing.
  • Build regression test suites for LLM and agent applications.
  • Establish automated testing and quality gates.
  • Implement CI/CD pipelines for AI/LLM applications.
  • Ensure reproducibility and reliability of AI experiments and production systems.

9. Cloud & AI Platforms

Experience with at least one major cloud platform is desirable:

  • Microsoft Azure
  • AWS
  • Google Cloud Platform (GCP)

Particularly valuable experience includes:

  • Azure OpenAI
  • Google Vertex AI
  • Amazon Bedrock
  • Amazon Bedrock AgentCore

Required Skills & Experience

  • 4+ years of professional software development experience.
  • Strong and demonstrable expertise in Python — mandatory.
  • Strong backend engineering fundamentals.
  • Experience building production-grade, scalable software systems.
  • Solid understanding of AI/LLM application architecture.
  • Hands-on experience with agent systems, LLM evaluation, or RAG.
  • Strong understanding of APIs, distributed systems, concurrency, and asynchronous processing.
  • Experience with testing, debugging, monitoring, and production troubleshooting.
  • Ability to work independently on technically complex engineering problems.

Highly Desirable Skills

  • LangGraph
  • OpenAI Agents SDK
  • MCP (Model Context Protocol)
  • A2A
  • MLflow
  • FastAPI
  • Docker & Kubernetes
  • Redis
  • Kafka
  • PostgreSQL
  • OpenTelemetry
  • Grafana & Prometheus
  • Azure OpenAI
  • Vertex AI
  • Amazon Bedrock / AgentCore

Senior Python Developer – AI Runtime & Evaluation Engineer

★
★
★
★
★

Ranking: 2637

Key Responsibilities

1. AI Agent Architectures

  • Design and develop modern AI agent architectures.
  • Work with planning, tool calling, function calling, memory, and context engineering.
  • Implement semantic routing and determine when to use workflows vs. autonomous agents.
  • Work with Model Context Protocol (MCP) and, ideally, Agent-to-Agent (A2A) architectures.
  • Build and maintain multi-agent orchestration systems.
  • Work with frameworks such as LangGraph and OpenAI Agents SDK.

2. Agent Runtime Engineering

  • Design and implement robust agent execution loops and runtimes.
  • Develop state management and orchestration mechanisms.
  • Implement retries, checkpointing, streaming, and human-in-the-loop workflows.
  • Build reliable tool-execution pipelines.
  • Handle context management and concurrency.
  • Implement comprehensive observability and tracing for agent execution.
  • Improve runtime reliability, scalability, latency, and fault tolerance.

3. AI Evaluation & Evals — Critical Responsibility

  • Design and implement offline and online evaluation frameworks for AI systems.
  • Develop meaningful benchmarks and golden datasets.
  • Generate and manage synthetic evaluation datasets.
  • Implement LLM-as-a-Judge and pairwise comparison methodologies.
  • Build automated regression-testing suites for LLM applications.
  • Establish experiment tracking and ensure reproducibility.
  • Analyze model and agent performance to continuously improve quality.
  • Work with MLflow for experiment tracking and evaluation workflows.

4. RAG Engineering

  • Design and optimize Retrieval-Augmented Generation (RAG) pipelines.
  • Work on chunking strategies, embeddings, and reranking.
  • Work with vector databases and retrieval systems.
  • Measure and improve retrieval quality.
  • Optimize context selection and context-window utilization.

5. LLM Engineering

  • Develop effective prompting strategies.
  • Implement reliable structured outputs and JSON Schema.
  • Build and integrate function-calling workflows.
  • Understand tokenization, context windows, and reasoning models.
  • Optimize LLM applications for cost and latency.
  • Implement caching strategies to improve performance and reduce inference costs.

6. Backend & High-Performance Engineering

  • Develop scalable and maintainable backend services using Python.
  • Design high-performance APIs and distributed backend systems.
  • Build production-grade services using technologies such as:
    • FastAPI
    • Docker
    • Kubernetes
    • Redis
    • Kafka
    • PostgreSQL
  • Optimize backend systems for throughput, concurrency, reliability, and low latency.

7. Observability & Monitoring

  • Implement comprehensive observability across AI and backend systems.
  • Work with:
    • OpenTelemetry
    • Distributed tracing
    • Metrics
    • Logs
    • Agent execution traces
    • Cost monitoring
    • Latency monitoring
  • Experience with Grafana and Prometheus is highly desirable.

8. Testing & Quality Engineering

  • Develop comprehensive unit and integration tests.
  • Implement AI-specific evaluation testing.
  • Build regression test suites for LLM and agent applications.
  • Establish automated testing and quality gates.
  • Implement CI/CD pipelines for AI/LLM applications.
  • Ensure reproducibility and reliability of AI experiments and production systems.

9. Cloud & AI Platforms

Experience with at least one major cloud platform is desirable:

  • Microsoft Azure
  • AWS
  • Google Cloud Platform (GCP)

Particularly valuable experience includes:

  • Azure OpenAI
  • Google Vertex AI
  • Amazon Bedrock
  • Amazon Bedrock AgentCore

Required Skills & Experience

  • 4+ years of professional software development experience.
  • Strong and demonstrable expertise in Python — mandatory.
  • Strong backend engineering fundamentals.
  • Experience building production-grade, scalable software systems.
  • Solid understanding of AI/LLM application architecture.
  • Hands-on experience with agent systems, LLM evaluation, or RAG.
  • Strong understanding of APIs, distributed systems, concurrency, and asynchronous processing.
  • Experience with testing, debugging, monitoring, and production troubleshooting.
  • Ability to work independently on technically complex engineering problems.

Highly Desirable Skills

  • LangGraph
  • OpenAI Agents SDK
  • MCP (Model Context Protocol)
  • A2A
  • MLflow
  • FastAPI
  • Docker & Kubernetes
  • Redis
  • Kafka
  • PostgreSQL
  • OpenTelemetry
  • Grafana & Prometheus
  • Azure OpenAI
  • Vertex AI
  • Amazon Bedrock / AgentCore

  • Equipo
  • Evaluador
  • Manager
  • Agencia
  • Cliente

GFT Cliente

Cliente

★
★
★
★
★

Comentarios: 0

Sandra Lobero

Agencia

★
★
★
★
★

Comentarios: 0

Paco Romero

Agencia

★
★
★
★
★

Comentarios: 0

Teba Gomez-Monche

Agencia

★
★
★
★
★

Comentarios: 0

claudia herrero

Agencia

★
★
★
★
★

Comentarios: 0

Hugo Herrero

Manager

★
★
★
★
★

Comentarios: 0

Víctor M. herrero

Evaluador

★
★
★
★
★

Comentarios: 3