Project details
Recommended projects
Evaluation Scenario Writer (m/w/d)
AI Evaluation Consultant (m/w/d)
Software Test Engineer (m/f/d)
Vibe Coding Web Scraping Expert (m/f/d)
System Engineer Functional Safety (m/f/d) / Functional Safety
Freelance Automotive Engineer (with Python) - Quality Assurance / AI Trainer
Senior Project Manager Customer Interaction
ERP Transformation Manager (m/f/d)
Freelance Product Owner for Point of Sale App
Freelance Mechanical Engineer with Python Experience (m/f/d)
Freelance Cybersecurity Consultant for AI Red Teaming
Commissioning & Qualification (C&Q) Engineer (m/w/d)
Senior Factor 10 Developer (IPS / IPM) (m/f/d)
AI Consultants - Data Science (m/w/d)
IT Project Manager ISO 27.001 - Gap Closure (m/f/d)
Interim Staff Product Manager (m/w/d)
AI Consultant - Machine Learning (m/w/d)
Management Consultant (Senior Level) (m/f/d)
HSE Specialist – Cell Manufacturing
Quality Compliance Auditor (GCP/GCLP/GVP) (M/W/D)
Senior Regulatory Compliance Expert (FDA Inspection Preparation) (m/f/d)
Tax Strategy Consulting
Java IT Architect (m/f/d)
Expert in process automation for law firm environments (m/f/d)
Kajabi Expert (m/f/d)
Safety and Health Protection Coordinator (SiGeKo) and Safety Specialist (SiFa) (m/f/d)
Control System Technician / Control Systems Specialist (m/f/d)
TM1 Planning Analytics and Interfaces Development (m/f/d)
Cyber Security Consultant – Product Security & Regulatory Compliance (m/f/d)
Senior Cloud Developer TypeScript (m/f/d)
Frontend developer to HR platform with Angular experience
Time's up! We are no longer accepting applications.
AI Agent Evaluation Analyst (m/w/d)
Project info
- Daily ratefrom 280€
- Language
- English(Advanced)
- English
- Remote100%
Description
We’re on the hunt for QAs for autonomous AI agents for a new project focused on validating and improving complex task structures, policy logic, and agent evaluation frameworks. Throughout the project, you’ll have to balance quality assurance, research, and logical problem-solving. This project opportunity is ideal for people who enjoy looking at systems holistically and thinking through scenarios, implications, and edge cases.
You do not need a coding background, but you must be curious, intellectually rigorous, and capable of evaluating the soundness and consistency of complex setups. If you’ve ever excelled in things like consulting, CHGK, Olympiads, case solving, or systems thinking — you might be a great fit.
What you’ll be doing:
- Reviewing evaluation tasks and scenarios for logic, completeness, and realism.
- Identifying inconsistencies, missing assumptions, or unclear decision points.
- Helping define clear expected behaviors (gold standards) for AI agents.
- Annotating cause-effect relationships, reasoning paths, and plausible alternatives.
- Thinking through complex systems and policies as a human would to ensure agents are tested properly.
- Working closely with QA, writers, or developers to suggest refinements or edge case coverage.
Requirements
- Excellent analytical thinking: Can reason about complex systems, scenarios, and logical implications.
- Strong attention to detail: Can spot contradictions, ambiguities, and vague requirements.
- Familiarity with structured data formats: Can read, not necessarily write JSON/YAML.
- Can assess scenarios holistically: What's missing, what’s unrealistic, what might break?
- Good communication and clear writing (in English) to document your findings.
We also value applicants who have:
- Experience with policy evaluation, logic puzzles, case studies, or structured scenario design.
- Background in consulting, academia, olympiads (e.g. logic/math/informatics), or research.
- Exposure to LLMs, prompt engineering, or AI-generated content.
- Familiarity with QA or test-case thinking (edge cases, failure modes, “what could go wrong”). Some understanding of how scoring or evaluation works in agent testing (precision, coverage, etc.).