You do not need a coding background, but you must be curious, intellectually rigorous, and capable of evaluating the soundness and consistency of complex setups. If you’ve ever excelled in things like consulting, CHGK, Olympiads, case solving, or systems thinking — you might be a great fit.

What you’ll be doing:

Reviewing evaluation tasks and scenarios for logic, completeness, and realism.
Identifying inconsistencies, missing assumptions, or unclear decision points.
Helping define clear expected behaviors (gold standards) for AI agents.
Annotating cause-effect relationships, reasoning paths, and plausible alternatives.
Thinking through complex systems and policies as a human would to ensure agents are tested properly.
Working closely with QA, writers, or developers to suggest refinements or edge case coverage.

Excellent analytical thinking: Can reason about complex systems, scenarios, and logical implications.
Strong attention to detail: Can spot contradictions, ambiguities, and vague requirements.
Familiarity with structured data formats: Can read, not necessarily write JSON/YAML.
Can assess scenarios holistically: What's missing, what’s unrealistic, what might break?
Good communication and clear writing (in English) to document your findings.

We also value applicants who have:

Experience with policy evaluation, logic puzzles, case studies, or structured scenario design.
Background in consulting, academia, olympiads (e.g. logic/math/informatics), or research.
Exposure to LLMs, prompt engineering, or AI-generated content.
Familiarity with QA or test-case thinking (edge cases, failure modes, “what could go wrong”). Some understanding of how scoring or evaluation works in agent testing (precision, coverage, etc.).

Project details

Recommended projects

AI Agent Evaluation Analyst (m/w/d)

MCP & Tools Python Developer (m/w/d)

Senior Data Architect (m/f/d)

Freelance Mechanical Engineer with Python Experience (m/f/d)

Freelance Cybersecurity Consultant for AI Red Teaming

AI Evaluation Consultant (m/w/d)

Freelance Electrical Engineer with Python Experience (m/w/d)

Freelance Automotive Engineer (with Python) - Quality Assurance / AI Trainer

Freelance Physics Expert (with Python) - Quality Assurance / AI Trainer

Freelance Java Developer (all genders)

Freelance Ruby Developer (m/f/d)

Freelance Biology Expert for AI Model Training (m/f/d)

Freelance Chemistry Expert for AI Model Training (m/f/d)

Evaluation Scenario Writer (m/w/d)

AI Consultant - Machine Learning (m/w/d)

AI Consultant for Vibe Coding (m/w/d)

Freelance Statistics Expert with Python Experience (m/f/d)

Freelance Civil Engineer with Python Experience (m/f/d)

Dentist for Training AI Models (m/f/d)

AI Consultants - Data Science (m/w/d)

Data Engineer (m/f/d)

SAP FI/CO Consultant (m/f/d) – Focus SAP R/3 - S/4HANA Transition

Mathematician with Python Experience (m/w/d)

Physicist with Python Experience (m/w/d)

Chemist with Python Experience (m/w/d)

Biologist with Python Experience (m/w/d)

Sales Manager for a Media Company (m/f/d)

Senior Regulatory Compliance Expert (FDA Inspection Preparation) (m/f/d)

Commissioning & Qualification (C&Q) Engineer (m/f/d)

Quality Compliance Auditor (GCP/GCLP/GVP) (M/W/D)

Frontend developer to HR platform with Angular experience

AI Agent Evaluation Analyst (m/w/d)

Project info

Description

Requirements