Get new jobs by email
- ...We are looking for an Agent Evaluation Engineer to design and build enterprise-grade evaluation frameworks for agentic AI systems. In this role, you will be responsible for creating automated evaluation pipelines, deployment quality gates, and reliability measurements that...Suggested
- ...We are looking for a Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. In this role, you will design evaluation frameworks, build custom evaluation pipelines, and establish automated...Suggested
- ...We are looking for a Python Engineer – Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI...SuggestedContract work
- ...We are looking for a Platform Engineer – CI/CD Gate & Online Evaluation to help establish automated quality controls for enterprise AI platforms and agent-based systems. In this role, you will integrate evaluation workflows into deployment pipelines, configure online quality...Suggested
- ...Eligible Countries Hong Kong S.A.R., Japan About the Opportunity We are seeking experienced Corporate/M&A Lawyers to help evaluate how effectively advanced AI systems answer real-world questions involving Japanese law . You will serve as an expert legal reviewer...Suggested40 hours per weekContract workFor contractorsRemote jobFlexible hours
- ...operating their application, ensuring a set of compliance and best practices. What you will do Design and implement automated evaluation frameworks for LangGraph-based agent workflows and orchestration pipelines. Develop build-time evaluation suites covering agent behavior...Suggested
- ...suites using Python and pytest. Validate permit/deny decisions, policy precedence rules, parameter-level access controls, and policy evaluation correctness. Test transitions between LOG_ONLY and ENFORCE policy modes and verify expected enforcement behavior. Perform API...Suggested
- ...Countries Japan, Hong Kong S.A.R. About the Opportunity We are seeking experienced litigation and disputes lawyers to help evaluate how effectively advanced AI systems answer real-world questions involving Japanese law . You will act as an expert legal...Suggested40 hours per weekContract workFor contractorsRemote jobFlexible hours
- ...Rate: $15/hour Hours: 4–20 hours per week About the Role We are looking for a Japanese-speaking AI Quality Analyst to evaluate personalized AI interactions. You will create realistic, multi-turn prompts based on your personal experiences and assess how effectively...SuggestedContract workRemote job
- ...equivalent) • Stateful distributed system patterns (checkpointing, idempotency) Will be a plus: • AWS AgentCore Runtime maturity evaluation experience • AWS Bedrock Memory API • Enterprise HITL approval workflow integration What it’s like to work at Intellias At...Suggested