





Remote, generic title and mid-level appeal create moderate applicant competition despite specialized infra requirements.
Platform and agent-evaluation skills transfer to other infra roles but AI-specific evaluation increases domain specificity.
Role has specific infra, sandboxing, and observability requirements without explicit years, implying medium strictness.
Login to See Your Match Score
Create a free account or log in to unlock your CV match score across:
Develop and maintain the core agent loop and execution model handling orchestration, concurrency, state management, and failure semantics for autonomous AI systems.
Build and oversee evaluation systems including task suites, verifiers, regression gates, and statistical methods to measure and improve agent performance.
Collaborate across engineering teams and research to design evaluation frameworks and interpret agent behavior post-training.
Strong software engineering experience with advanced Python; demonstrated ability to ship and operate production services.
Experience with concurrency and distributed systems including async execution, worker pools, retries, and cancellation handling.
Proficiency in containers and sandboxing technologies such as Docker and OCI internals for reproducible, isolated environments.
Work Experience Required: Not explicitly mentioned in the JD.
Experienced in building complex systems involving concurrency, state management, and measurement infrastructure rather than only AI prompting.
Comfortable owning loosely specified problems with systematic debugging and strong observability skills (distributed tracing, structured logging).
Has background or interest in evaluation/benchmarking for AI systems, agentic systems, LLM inference, or developer tooling.