In brief

Yes. The AI Proof Work checker can help you prioritize one recurring task for investigation after an agent benchmark result. First check what the benchmark tested and whether its task resembles your work. The checker provides task-level change-pressure signals; it does not validate the benchmark, measure employer adoption, or estimate your chance of job loss. Use local workflow evidence before changing how you work or making a career decision.

What should I do after an agent benchmark result?

Yes: use the AI Proof Work checker to help prioritize one recurring task for investigation, but first establish what the benchmark tested and whether that test resembles your work. A benchmark result is evidence about a specified system performing on a selected task set. METR’s “Task-Completion Time Horizons of Frontier AI Models” estimates success thresholds from human task durations on a mainly technical suite. METR also says those tasks are self-contained and well-specified, unlike day-to-day work dependent on prior context. So a headline result may make a task worth examining sooner; it does not establish that an agent can perform your whole role. The checker adds a different lens: it provides task-level change-pressure signals for work you describe. It does not validate or reproduce a third-party benchmark, measure employer adoption, or calculate your chance of losing a job. Treat the sequence as benchmark result, plausible match to one recurring task, checker-assisted prioritization, then proportionate local investigation. Before changing a workflow or paying for training, consider the task’s inputs and context, expected output, failure cost, review burden, and accountable decision-maker. A high result is a reason to investigate a close match, not to abandon a role. If the match is weak, choose another task or stop extrapolating. The evidence needed for a career decision is broader than either signal.

Sources: Task-Completion Time Horizons of Frontier AI Models; AI Job Risk Checker

What does the benchmark actually measure?

A benchmark reports how a specified model-agent setup performed on a selected task set under a particular scoring and repetition method. Its result belongs to that setup and task distribution, not automatically to a job or worker. To interpret it, identify the model and version, scaffold and tools, task selection, instructions, success criteria, number of attempts, aggregation method, uncertainty, and update date. Each choice narrows the claim the result can support.

METR’s [Task-Completion Time Horizons of Frontier AI Models](https://metr.org/time-horizons/) shows why the metric definition matters. METR estimates human-expert completion time for each task, then fits a logistic curve predicting an agent’s success probability as task duration changes. The 50% time horizon is the human task duration where the fitted curve predicts 50% success; the 80% horizon uses an 80% threshold. This is a modeled threshold over benchmark tasks, not how long an agent can operate autonomously, a promise of success on every task of that duration, or a measure of a worker’s day. The current suite draws on RE-Bench, HCAST, and shorter software tasks, mainly in software engineering, machine learning, and cybersecurity. METR describes these tasks as self-contained, well specified, and automatically evaluated. That scope defines the result and limits generalization to other work.

The human-time baseline matters. METR says that for most tasks it contracts people to attempt them and uses the geometric mean of successful completion times; participants generally receive the same instructions and affordances as agents. For some tasks, it uses expert estimates or QA completion times because reliable human timings are unavailable. The dashboard warns that estimates above 16 hours are unreliable with the current suite. A headline omitting the baseline, task mix, or limit can sound broader than the measurement. An eight-hour horizon means a fitted threshold against human task duration within this suite; it does not mean eight hours of ordinary professional work can be delegated end to end.

Version history matters. In [Time Horizon 1.1](https://metr.org/blog/2026-1-29-time-horizon-1-1/), METR reports expanding its suite from 170 to 228 tasks, adding 73, removing 15, and updating 53; it re-estimated 14 of 33 previously measured models. The maintainer attributes estimate changes mainly to suite changes and run noise, and says confidence intervals remain wide. Only five of 31 tasks estimated at eight hours or more had measured human baselines. This does not make the benchmark useless; it means comparisons across releases involve measurement changes as well as model changes. Before mapping a result to your work, note the exact release and setup, then state the narrow capability claim the task set supports. The checker can help prioritize a task for further investigation; it cannot fill missing benchmark details or turn capability into workplace evidence.

Sources: Task-Completion Time Horizons of Frontier AI Models; Time Horizon 1.1

How close is the benchmark task to one task in my work?

Compare five properties before treating a benchmark task as a candidate for your own work: its inputs and context; the tools and environment available; the expected output and scoring rule; the errors it allows or tests; and the human review or accountability around completion. This is a practical transfer test, not a validated scale. Similar labels such as “coding,” “research,” or “customer support,” and verbs such as “write” or “summarize,” establish little by themselves. The question is whether the conditions that make the benchmark task succeed also occur in a recurring task you perform. METR describes its time-horizon tasks as software engineering, machine learning, and cybersecurity. They are designed to be self-contained, well-specified, and clearly scored, often automatically. The page also explains that its task-duration estimate uses human expert completion time and a fitted success curve. METR cautions that professionals doing equivalent work may have more project context than its evaluators and agents receive. That detail matters for transfer: a measured task can be technically demanding while omitting the history, conventions, and coordination that shape the same kind of work inside an organization. Imagine a software worker sees a benchmark result for an agent asked to fix a bug in a supplied code repository. The first property is inputs: does the supplied repository include the relevant design decisions, issue history, dependencies, and undocumented conventions, or only the files needed for a bounded fix? Second, environment: can the agent run tests and inspect the same tools and permissions available in the recurring task, or is the benchmark environment specially prepared? Third, output and scoring: does passing a test mean the change is maintainable, compatible with release requirements, and acceptable to the team, or only that a defined test condition passed? Fourth, errors: are failures reversible in a sandbox, or could a mistaken change affect users, security, or production data? Fifth, review: who checks the patch, resolves ambiguous requirements, and owns the release decision? Some properties may transfer. If the worker regularly handles isolated, well-described defects in a familiar repository, with reproducible tests and a reviewer who can verify proposed changes, a benchmark task may resemble one step. Other properties may not. A recurring assignment can depend on prior project choices, tacit conventions, changing requirements, clarification from colleagues, and accountability for a release. Those differences do not prove an agent cannot help; they identify what the benchmark did not establish. A pass on a self-contained task is evidence about that tested setup and task, not proof of independent performance across the larger workflow. A broader workflow suite can help reveal dimensions to inspect. OSWorld 2.0 describes 108 long-horizon computer-use workflows and calls out challenges such as cross-source reasoning, implicit state, and dynamic environments. Its task design is a useful contrast to short isolated tests, but the suite remains a set of constructed workflows with its own systems, tasks, and scoring. “More work-like” does not mean identical to your employer’s process. Use the five properties to identify the closest recurring task, then investigate that task; if the match is weak, do not stretch a benchmark result into an occupation-wide conclusion.

Sources: Task-Completion Time Horizons of Frontier AI Models; OSWorld 2.0

Which details in a benchmark report change how much weight it deserves?

A benchmark result deserves more weight when the tested agent, task conditions and success standard resemble the work you want to understand, and when the report shows result stability. A changed suite, sparse human baselines, unclear scoring or an older setup should narrow the conclusion. These are reasons to read carefully, not dismiss a benchmark. Record its release and configuration, then state what the result supports.

METR’s January 2026 “Time Horizon 1.1” update shows why release details matter. METR expanded its suite from 170 to 228 tasks, adding 73, removing 15 and updating 53; definitions, human time estimates and scoring functions changed. It moved evaluation infrastructure from Vivaria to Inspect. Of 33 models previously estimated under TH1, 14 were re-estimated under TH1.1. METR attributes estimate changes mainly to suite changes and run variation, with a smaller contribution from infrastructure. Figures across releases are not repeat measurements under identical conditions. This does not make either version worthless; it makes a cross-release difference partly a measurement-design question, not automatically a capability change.

The update makes uncertainty concrete. The larger suite tightened estimates, especially at the upper end, but METR says confidence intervals remain wide. Among 31 tasks estimated to take humans at least eight hours, only five had measured human baselines; the rest used estimates. Full-trend comparisons between TH1 and TH1.1 cannot be direct because no pre-2023 models were re-estimated with TH1.1. This is uncertainty within the benchmark: dependence on sampled tasks, baselines, runs and release design. A point estimate can look crisp while meaningful uncertainty remains.

Workplace transfer is a different uncertainty. METR’s “Task-Completion Time Horizons of Frontier AI Models” defines a time horizon as a fitted success threshold based on human expert task duration, not agent operating time. Current measurements use over a hundred software tasks, mainly in software engineering, machine learning and cybersecurity. METR describes these as self-contained, well-specified tasks with clear, automatically evaluated criteria; estimates above 16 hours are unreliable with the current suite. Even a precise result cannot show whether the system handles your organization’s context, tools, exceptions, review or accountability. Transfer uncertainty asks whether measured capability applies to your recurring task; measurement uncertainty does not settle it.

Keep the report’s release, model and agent setup, tools or scaffold, task mix, scoring, runs and uncertainty beside the headline. Ask whether those conditions resemble your task and its quality bar. If suite or setup changed, compare like with like where possible; avoid treating trends as more exact than reported. The bounded conclusion: the benchmark provides evidence about performance on defined tasks under a stated setup. Closer task fit and clearer reliability give more reason to investigate corresponding work. Neither establishes employer adoption or predicts what happens to your job.

Sources: Time Horizon 1.1; Task-Completion Time Horizons of Frontier AI Models

Can a benchmark score be misleading even when its task sounds relevant?

Yes. A benchmark task can resemble part of your work and still produce a misleading score if its instructions, tests, or scoring rule do not measure completion. Task relevance and benchmark quality are separate checks: the first asks whether tested work resembles your task; the second asks whether the evaluation rewards and penalizes the right behavior. A result deserves less weight when the test cannot reliably distinguish a valid solution from a flawed one.

A July 2026 audit of SWE-Bench Pro by OpenAI illustrates this issue in one coding benchmark. On its 731-task public split, OpenAI reported frontier-model pass rates rising from 23.3% to 80.3% over eight months. The later audit found quality problems. An initial pipeline flagged 286 potentially broken tasks for deeper review. Within that flagged subset, the pipeline identified 200 as broken, or 27.4% of the public split; a separate human annotation campaign identified 249, or 34.1%. Five experienced software engineers reviewed each task in that human-reviewed subset. The article estimates about 30% of SWE-Bench Pro tasks are broken. The 200 and 249 counts come from distinct review pathways and should not be added together or treated as a universal defect rate for agent benchmarks.

The audit describes four ways scores can diverge from useful work. An overly strict test may reject a functionally correct solution for an implementation detail the prompt never required. An underspecified prompt may omit requirements that hidden tests enforce and that a solver could not reasonably infer. Low-coverage tests can let an incomplete fix pass. A misleading prompt can direct a system toward behavior that conflicts with the tests. The first two can create false failures; the latter two can create false passes. Depending on the flaw, a benchmark may understate capability or overstate task completion.

This audit is a reason to inspect evidence, not to discard benchmarks. Its review can help improve evaluation by making flaws visible. The audit concerns one dataset and comes from a model developer with a stake in coding evaluations, so its estimate should remain specific to SWE-Bench Pro. It does not show that another suite is flawed or that benchmark results have no value.

When a result sounds relevant to your task, inspect prompt examples, success criteria, test coverage, audit notes, and whether passing means meeting your actual quality bar. Ask whether a failure reflects unclear instructions, or a pass leaves requirements unchecked. If the report does not expose these details, narrow the conclusion you draw from its headline. Relevance makes a benchmark worth investigating; trustworthy scoring determines how much weight its result can carry.

Sources: Separating Signal from Noise in Coding Evaluations

What does a more work-like benchmark add—and what does it still leave out?

A workflow benchmark can reveal problems that short, isolated tests may miss: tracking information across sources, responding to a changing environment, inferring unstated context, and checking whether a multi-step result is complete. But “more realistic” still describes a sample built for evaluation. It can sharpen the question you ask about a task; it cannot tell you whether your employer has adopted an agent or whether your role will change. OSWorld 2.0 illustrates what broader workflow design can add. Its paper describes 108 long-horizon computer-use workflows across everyday and professional tasks. They target dynamic environments, cross-source reasoning, implicit-state inference, and visual-spatial precision. Under its primary binary-completion measure at 500 steps, the paper reports that its best-performing setup completed 20.6% of tasks and achieved a 54.8% partial score. Those are results for this benchmark, model configuration, and scoring setup, not a general estimate of computer work agents can perform. Reported failures describe mechanisms: agents may lose track of constraints, miss information that arrives mid-task, guess rather than ask, or skip verification. If a task in your work depends on one of these abilities, inspect the match. It is not enough that both the benchmark and your task involve “computer use.” Ask whether they share inputs, state changes, tools, acceptable outputs, and ways to recover from errors. METR’s Task-Completion Time Horizons of Frontier AI Models shows why a longer test is not automatically a closer work match. METR says its suite consists primarily of software engineering, machine learning, and cybersecurity tasks. These are designed to be self-contained and well-specified, with clear criteria that can be evaluated automatically. METR cautions that human-duration estimates may omit context professionals bring to day-to-day work. OSWorld 2.0 samples different workflow mechanics; METR measures a different defined task set. Neither label alone says whether a suite resembles a particular worker’s task. “More work-like” is relative to the evaluation being compared. A constructed environment can include realistic files or state while leaving out team handoffs, permissions, unusual exceptions, customer expectations, and responsibility for consequential decisions. Benchmarks make performance comparable under bounded conditions; that strength also limits what they establish. Prefer the benchmark whose task mechanics most closely match the work you want to investigate, not simply the newest or most realistic-sounding one. Use the AI Proof Work checker to help prioritize a task you describe, then, where policy permits, compare a supervised trial or observation with the current workflow. Check quality, corrections, review time, context handling, recovery, and accountability. A benchmark can point to the right question; evidence from the actual task should decide what to do next.

Sources: OSWorld 2.0; Task-Completion Time Horizons of Frontier AI Models

Why does a task-level checker answer a different question from a benchmark?

A benchmark and a task checker can both help narrow attention, but they organize different evidence. A benchmark asks whether a defined agent setup completed selected tasks under a stated scoring rule. The AI Proof Work checker, as described for this publication, gives task-level change-pressure signals for work a reader describes. It can help choose a task to examine; the product information does not establish that it imports, verifies, or reproduces a third-party benchmark result. Treating the two outputs as comparable scores would blur what each measures. The distinction matters because a benchmark result is capability evidence within its test conditions. It says something about a system, its instructions, tools, task selection, and evaluation method. It does not show that an employer has deployed that system, redesigned a role around it, or reduced staffing. A checker signal has a different boundary: it helps frame which recurring activities may merit task-level investigation. It is neither proof that a task is already being automated nor a probability that a worker will lose a job. The ILO–NASK study, “Generative AI and Jobs: A Refined Global Index of Occupational Exposure,” illustrates why task-based exposure is useful while also showing its limits. The 2025 index combined nearly 30,000 occupational tasks with expert validation, AI-assisted scoring, and harmonized labor data to estimate potential exposure to generative AI across occupations. The ILO’s summary explicitly says these figures describe potential exposure, not actual job losses, and that implementation depends on technological constraints, infrastructure, skills, and other conditions. This is a broad occupational analysis, not a report on agent benchmark performance or a specific workplace. Observed use and employer adoption require their own evidence. A worker might try an approved tool on a task without the organization adopting it as standard practice. A team might introduce a system but retain human review, alter handoffs, or use it only for a narrow step. To establish local adoption, look for evidence in the relevant workflow: an authorized process, repeated use, changed responsibilities, or a formal deployment decision. A benchmark leaderboard cannot supply that evidence, and a task checker does not claim to. A practitioner resource, “AI Agent Benchmark,” frames evaluation around real workflows, expected outputs, failure risks, and whether work can be trusted and checked. That is useful context for why benchmark-to-work transfer deserves scrutiny, but it does not validate this checker or establish employment effects. Use the checker after identifying a plausible recurring task match, then investigate the task’s actual conditions. The benchmark can prompt the question; the checker can help prioritize where to look. Neither is a second adoption signal. A task-level change-pressure result is a starting point for inquiry, not a verdict about a job or workplace.

Sources: AI Job Risk Checker; Generative AI and Jobs: A Refined Global Index of Occupational Exposure; AI Agent Benchmark

What evidence should I gather around the task before deciding?

Before deciding what a benchmark result means for your work, make a short record of one recurring task as it happens now. Include how often it comes up, what information and prior context it needs, what counts as an acceptable output, the steps required to complete and review it, and what happens if an error gets through. Note how the work is recovered when something goes wrong, whether information may be entered into an approved tool, and who is accountable for the final decision.

Then keep four evidence types separate. The benchmark report describes tested capability: its model and agent setup, task suite, instructions, scoring rule, and limits. For example, METR’s “Task-Completion Time Horizons of Frontier AI Models” says its current suite is mainly software engineering, machine learning, and cybersecurity work; tasks are designed to be self-contained, with clear criteria that can be automatically evaluated. METR also cautions that day-to-day professional work often draws on prior conversations, tacit knowledge, or familiarity with a codebase. It describes tested conditions, not whether a system can perform your whole recurring task. OSWorld 2.0 describes 108 computer-use workflows, including challenges such as cross-source reasoning, changing environments, and hidden state. Those features can prompt useful questions about your task, but the constructed suite cannot show whether your organization uses an agent or accepts its outputs.

The checker supplies another lens: task-level change-pressure signals based on work you describe. It does not reproduce or verify the benchmark, and its signal is neither employer-adoption evidence nor a displacement probability. The ILO–NASK “Generative AI and Jobs: A Refined Global Index of Occupational Exposure” combines nearly 30,000 occupational tasks with expert validation, AI-assisted scoring, and harmonized labor data. It measures potential exposure, not a particular team’s implementation. Use the checker to choose a task to examine, not to confirm benchmark transfer or workplace change.

For practical fit, gather permitted evidence from the workflow itself. Where policy allows, compare the current process with assistance through a bounded supervised trial or structured observation. Record whether the output meets the agreed standard, what needs correction, time spent reviewing, context the system misses, and how errors are detected and recovered. Include handoffs and approval work; a faster first draft may not shorten the whole task if review or rework grows. Do not place restricted work data in an unapproved service just to test a hypothesis. When a trial is not permitted, use approved documentation, policy, process changes, or direct clarification about whether the workflow is being evaluated or adopted. A low-stakes task with clear checks may warrant a small supervised test; a sensitive task, high-cost error, unclear recovery path, or missing approval calls first for policy and accountability clarification.

Sources: Task-Completion Time Horizons of Frontier AI Models; OSWorld 2.0; Generative AI and Jobs: A Refined Global Index of Occupational Exposure

Should I change my workflow, learn a tool, or make a bigger career move?

Match the size of your response to the evidence you have. A benchmark result justifies investigating a plausible task match, not replacing a workflow or leaving a field. A permitted, bounded trial can test whether assistance improves one step. Repeated local evidence may justify learning a tool or discussing changed responsibilities. A larger move needs broader evidence and must fit your needs. This sequence keeps the question practical: what evidence would change the next decision, and what is still unknown? Task exposure is not an employment outcome. The International Labour Organization’s “Generative AI and Jobs: A Refined Global Index of Occupational Exposure” combines nearly 30,000 occupational tasks with expert validation, AI-assisted scoring, and harmonized labor data. Its global estimate concerns potential exposure; the ILO explicitly says it is not a count of actual job losses and notes that implementation depends on conditions such as infrastructure and skills. The report says transformation is more likely than full replacement. This cautions against turning benchmark capability into a personal forecast; it says nothing about your employer’s actions or your role’s future. For a task with a clear quality check and limited error consequences, observe the process or, if rules allow, run a supervised trial on approved material. Compare output quality, corrections, review time, handoffs, and remaining human decisions. A faster first draft is not enough if checking and repair absorb the apparent saving. Keep sensitive work data out of tools unless your organization permits that use. For hard-to-verify or consequential tasks, clarify policy and accountability first. If observations show a workflow change, consider upgrading your current role before a wholesale pivot. Learn what supports that workflow: task framing, context, output checks, and when to stop using the tool. A course can teach a specific skill; a project can demonstrate its application. Neither alone proves hiring value or readiness for another profession. Choose a degree when its depth or credential serves the goal, not as a default reaction to a benchmark. Consider an adjacent move only if evidence points to a durable shift in your task mix or local demand, and the move uses experience you already have. Check prerequisites, salary floor, location, training time and cost, health, and family responsibilities. The free AI Proof Work checker can help prioritize a task you describe using task-level change-pressure signals; it does not validate the benchmark or estimate your chance of job loss. For constrained comparisons among staying, adjacent moves, and larger changes, the career roadmap may help after evidence gathering. Prefer the smallest reversible learning or workflow step that answers the open question.

Sources: AI Job Risk Checker; Generative AI and Jobs: A Refined Global Index of Occupational Exposure

What is one realistic next step after seeing the result?

After an agent benchmark result, write down three recurring tasks from your week. Choose the one whose inputs, tools, expected output, and success test most resemble the benchmark. Record its context, quality bar, review burden, and who is accountable. This comparison does not show that your workplace uses the benchmarked system.

The METR time-horizon documentation shows why the match matters: its results concern a defined collection of mostly self-contained technical tasks with clear evaluation criteria. If none of your tasks resembles that setup, stop; do not stretch a headline result across your job. If one does, the free AI Proof Work checker can help prioritize it using task-level change-pressure signals. It does not validate the benchmark or turn either signal into a personal displacement probability.

Gather evidence proportionate to the decision. Where policy permits, observe or run a supervised trial and compare the full task with the current process: quality, context handling, correction time, error recovery, and accountable review. Keep restricted data out of unapproved tools. If no safe trial is available, look for local workflow or implementation evidence before changing practice. A closer benchmark match, reliable local results, or confirmed implementation could change what deserves attention. Until then, investigate one task and tie larger training or career decisions to your constraints and stronger evidence.

Sources: AI Job Risk Checker; Task-Completion Time Horizons of Frontier AI Models

Questions readers ask

Does the AI Proof Work checker validate an agent benchmark?

No. The checker provides task-level change-pressure signals for work you describe. The available product information does not establish that it imports, validates, or reproduces third-party benchmark results.

Does a strong agent benchmark result mean my job is at risk?

No. A benchmark result describes a specified system performing on a selected task set under particular conditions. It does not establish employer adoption or predict an individual’s chance of displacement.

What should I compare before applying a benchmark to my work?

Compare the benchmark and your recurring task across inputs and context, tools and environment, expected output and scoring, error conditions, and human review or accountability. Then gather proportionate evidence from the local workflow.

Sources and notes

  1. Task-Completion Time Horizons of Frontier AI Models

    METR documents its time-horizon method, task scope, success thresholds, and limits on transfer to other work.

  2. Time Horizon 1.1

    METR reports its TH1.1 task-suite revisions, model re-estimation, and continuing uncertainty.

  3. OSWorld 2.0

    The paper describes 108 computer-use workflows and reports results under its stated benchmark setup and scoring.

  4. Separating Signal from Noise in Coding Evaluations

    The SWE-Bench Pro audit describes task and scoring issues in that benchmark and its review methods.

  5. Generative AI and Jobs: A Refined Global Index of Occupational Exposure

    The ILO summary describes a task-based exposure method and distinguishes potential exposure from observed job losses.

  6. AI Job Risk Checker

    The saved publication assignment defines the checker as providing task-level change-pressure signals, not displacement probability.

  7. AI Agent Benchmark

    The opened practitioner resource frames agent evaluation around workflows, outputs, failure risks, and checking work.

Apply this to your own work

See the whole job market at once.

Explore which occupations AI may reshape, then turn the signal into a practical response.

Explore the job map