Build realistic tasks
We collect agent tasks that require planning, tool use, information synthesis, and artifact delivery across realistic desktop and web environments.
Agent Evaluation Suite
AIPERT builds benchmark tasks, scoring rubrics, and inspection tools for AI agents that operate across files, browsers, code, documents, data, and long-horizon workflows.

Benchmark Design
We collect agent tasks that require planning, tool use, information synthesis, and artifact delivery across realistic desktop and web environments.
Each task includes expected outputs, constraints, validation hooks, and review notes so performance can be inspected instead of guessed.
Agents are evaluated with execution traces, intermediate artifacts, final deliverables, and expert review where pure automation is not reliable enough.
Evaluation Philosophy
AIPERT focuses on whether an agent can finish work under constraints: read source material, operate tools, revise artifacts, recover from errors, and leave an auditable record of what happened.
Scenarios preserve the messy structure of real work while keeping evaluation boundaries explicit.
Scoring distinguishes exact checks, partial credit, unsafe behavior, and unresolved review items.
Evaluation includes the path an agent took, not only the final answer it produced.
Public claims will be tied to documented task suites, versioned cases, and reproducible evidence.
Task Domains
Repository navigation, issue fixing, regression testing, documentation, and code review.
Spreadsheet analysis, web research, citation-aware synthesis, and structured extraction.
Slides, reports, PDFs, spreadsheet artifacts, formatting QA, and revision tasks.
Multi-step browsing, form interaction, page inspection, comparison, and evidence capture.
Command-line work, local files, service checks, deployment routines, and recovery from errors.
Instruction following, boundary handling, risk awareness, and refusal quality under pressure.
Partners & Contributors
AIPERT's collaboration network will include university labs, workflow experts, evaluation engineers, and student contributors. Confirmed institutions and collaborators will be published here.
Agent evaluation methodology, task construction, and result analysis.
Real workflows, industry constraints, and expert review standards.
Environment packaging, automated checks, trace capture, and regression validation.
Task curation, annotation review, documentation, and case maintenance.
Roadmap
Benchmark overview, evaluation principles, and representative task categories.
Versioned evaluation cases with rubrics, environment notes, and validation artifacts.
Model and agent results with trace-backed analysis, limitations, and review notes.
Agent evaluation should make progress visible: what was attempted, what worked, what failed, and what still needs human review.
Resources
Contact
AIPERT welcomes conversations around task design, evaluation methodology, trace analysis, and domain-specific benchmark construction.
contact@aipert.top