In brief

When an AI tool suggests U.S. ICD-10-CM diagnosis codes, a qualified coder still needs to check that every proposed condition is supported by the complete record, the code reflects documented certainty and specificity, applicable conventions and sequencing are followed, and the code set is valid for the encounter date. Missing, conflicting or unclear clinical documentation needs provider clarification; a model cannot supply the clinical fact. AI can help find candidates, but the available evidence does not establish that a suggested code set is ready for autonomous claim submission.

Which coding judgments cannot be settled by a suggested code?

Code lookup is only one part of diagnosis coding. Under the FY2026 ICD-10-CM Official Guidelines, the entire record should be reviewed to establish the reason for the encounter and conditions treated. The guidelines also say accurate coding depends on consistent, complete documentation and describe a joint effort between provider and coder. That means a suggested label is a lead to verify, not evidence that the condition belongs on the claim. The FY2026 rules apply October 1, 2025 through September 30, 2026. [CMS’s official guidelines](https://www.cms.gov/files/document/fy-2026-icd-10-cm-coding-guidelines.pdf) support this boundary; they do not assess any AI product. (Sources 1)

Review whether the documentation supports the diagnosis at all, whether the code is as specific as the record permits, and whether the certainty expressed in the note allows reporting that condition in the setting at hand. The guidelines direct coders to use the highest specificity supported by the documentation, while also allowing symptom, sign or unspecified codes when those best reflect what is known. More specificity is not automatically better. For example, imagine a note names a chronic condition but does not document a stage. A system may produce a stage-specific candidate from surrounding text or common patterns; the coder must verify that stage is actually documented and applicable rather than infer it.

Then check whether the whole proposed set makes sense together. A code can look plausible on its own while missing a related condition, duplicating a concept, or conflicting with an instruction about sequencing or additional codes. The Alphabetic Index and Tabular List conventions take precedence over general guidelines, and the official guidance says all relevant sections must be considered. This makes code-set review a different task from matching one phrase to one code.

The boundary is sharper when records disagree or leave a clinical question unresolved. CMS guidance says missing, conflicting or unclear documentation should be resolved by the provider. The coder can identify the gap and route a query through the organization’s process, but should not turn a likely clinical explanation into a documented diagnosis. Human review matters most when the suggestion adds an unsupported condition, exceeds documented detail, changes reporting meaning, or should trigger clarification. It is a documentation and rule-application task, not simply a test of whether the code exists.

Sources: FY 2026 ICD-10-CM Official Guidelines for Coding and Reporting

What does the evidence say about suggestions versus autonomous coding?

The evidence supports treating AI output as a candidate list for review, not proof of a complete and correct claim. The U.S. Government Accountability Office’s July 2026 technology spotlight distinguishes tools that suggest codes for coder review from newer systems that may assign codes autonomously. GAO reports provider adoption but also says accuracy can be difficult to verify and few independent studies evaluate these tools. It notes possible patient harm and over- or under-reimbursement from inaccuracies. This is a policy and evidence overview, not a controlled comparison of vendors or evidence that all deployments perform alike. [GAO’s spotlight](https://www.gao.gov/products/gao-26-109116) therefore supports a need for measured oversight, not a universal error rate. (Sources 2)

Two research examples show why task and input conditions matter. A study indexed in PubMed compared six language models with one human coder on 50 deidentified inpatient notes. The human coder identified 165 unique codes; model outputs ranged from 221 to 658. That difference is a reason to inspect extra candidates and verify support, but the count alone does not prove which codes were wrong. The sample was small, used one coder as its reference, and does not represent every current system or production workflow. [The study abstract](https://pubmed.ncbi.nlm.nih.gov/39608761/) reports those limits and results. (Sources 3)

A separate 2024 JAMIA experiment used GPT-3.5-turbo-0613 in a study of ICD-10 data augmentation and direct coding. The authors generated 9,606 synthetic discharge summaries from code-description label sets, then evaluated direct coding separately on real MIMIC-IV notes. In the real-data test, GPT-3.5 scored 14.76% F1 under leaf-only scoring and 17.13% under set-based scoring; its scores on prompted synthetic text were higher. The authors concluded that GPT-3.5 alone in their tested prompt setting was unsuitable for ICD-10 coding. These results describe one older model, dataset and experiment, not a current product’s error rate. [The University of Edinburgh’s publisher-PDF copy](https://www.pure.ed.ac.uk/ws/portalfiles/portal/476341828/FalisEtalJAMIA2024CanGPT-3.5Generate.pdf) reports the study design, test results and conclusion. (Sources 4)

Together, these sources favor coder-reviewed use when diagnosis coding affects reporting and reimbursement. They do not prove that every routine candidate requires equal manual effort, or that review always catches every error. They do show why a benchmark, a vendor’s own reported result or a tool’s ability to retrieve a plausible code cannot substitute for checking the record, the complete code set and local audit findings.

Sources: Science & Tech Spotlight: AI for Medical Notes and Coding; Extracting International Classification of Diseases Codes from Clinical Documentation Using Large Language Models; Can GPT-3.5 generate and code discharge summaries?

What should a practical human review check first?

A useful claim audit starts with the record, not the code label. For every proposed code, locate the supporting documentation and confirm that it belongs to the encounter being coded. Look for additions that have no support and for relevant conditions that the tool may have missed. Then check that specificity and certainty match the documentation, and review the applicable Index, Tabular List instructions, sequencing and combination rules under the correct care setting. This is a practical synthesis of official guidance, not a newly issued regulatory checklist. [The FY2026 guidelines](https://www.cms.gov/files/document/fy-2026-icd-10-cm-coding-guidelines.pdf) establish the record-review and rule requirements. (Sources 1)

Next, verify code validity against the service or discharge date. CMS maintains separate fiscal-year code files and guidelines. Its ICD-10 page lists FY2027 ICD-10-CM files for dates beginning October 1, 2026, while the current FY2026 code files are organized by effective periods, including an April 2026 update. A system can return a recognizable code and still require confirmation that the correct version applies to this encounter. [CMS’s code-set index](https://www.cms.gov/medicare/coding-billing/ICD-10-codes) supports checking the effective release; it does not show whether a vendor updates its software on time. (Sources 5)

Finally, route ambiguity rather than resolving it by model inference. If the note conflicts with another part of the record, or a detail needed for a more specific code is absent, follow the provider-query and escalation process in your setting. Ask whether the tool preserves the source passage behind each suggestion, makes omissions and changes visible, records overrides, and allows review of the applicable code-set version. GAO’s concern that independent performance can be difficult to verify makes those workflow controls relevant, though the spotlight does not prescribe one universal design. (Sources 1, 2, 5)

For a bounded local check, ask a supervisor or compliance lead to approve a small sample of routine and exception cases. Record where review finds unsupported specificity, missing or extra codes, documentation conflicts, or stale code references. Use only an approved system for protected health information. This exercise can inform how much review different queues need in that organization; it cannot establish a general safety threshold from a handful of cases.

Sources: FY 2026 ICD-10-CM Official Guidelines for Coding and Reporting; Science & Tech Spotlight: AI for Medical Notes and Coding; ICD-10 | CMS

What is a realistic next move for a medical coder?

Build the parts of coding work that make a decision defensible: reading across the record, applying conventions, recognizing certainty and specificity limits, identifying when a query is needed, and explaining why a suggested code does not fit. Those foundations can transfer across tools. Learning one vendor interface may help with a current workflow, but the interface can change; it is a narrower investment than sharpening documentation interpretation and audit skills.

Compare three options against your real constraints. First, upgrade in place by learning the approved tool and auditing its outputs, especially the exceptions your queue sees. Second, investigate adjacent responsibilities such as coding quality, documentation integrity or compliance if those duties fit your experience and openings exist where you can work. As a broad labor-market signal, BLS projects 8% employment growth and about 14,000 annual openings for medical records specialists over 2025–35. Its Employment Projections program develops occupational employment estimates, then estimates openings by combining projected employment change with occupational separations—workers moving to other occupations or leaving the labor force. Many projected openings for this occupation are expected to come from replacement needs. This is a national projection for the whole occupation, which includes more than the adjacent duties named here; it does not establish local vacancies, hiring requirements or AI’s effect on a specific role. National projections and local job postings answer different questions, and neither predicts your outcome. [BLS’s occupation outlook](https://www.bls.gov/ooh/healthcare/medical-records-and-health-information-technicians.htm) reports the population and result; [its projections-method overview](https://www.bls.gov/emp/documentation/projections-methods.htm) and [occupational-separations method](https://www.bls.gov/emp/documentation/separations.htm) explain the basis for openings. Third, consider a larger career change only after checking salary needs, location, prerequisites, training cost and time, health and family responsibilities. These are context signals, not a guarantee of opportunity. (Sources 6, 7)

The supported verdict is to improve review and workflow capability before making a costly pivot. Start with one conversation: ask your supervisor which code families or exceptions remain human-reviewed, what errors and overrides are tracked, what evidence accompanies a suggestion, and whether learning time is paid. A new study of a current system showing reliable performance on your own representative records, with transparent auditing and a defined escalation path, could justify a different local allocation of review effort. Without that evidence, candidate-finding assistance and autonomous claim readiness remain separate claims.

If you want to map which parts of your own work are more exposed to task change, AI Proof Work’s [free task checker](/ai-job-risk-checker) provides change-pressure signals, not a probability of redundancy. If you need to compare staying and redesigning, adjacent options and a larger change against your pay floor, location, training time and family constraints, the [career roadmap](/career-roadmap) can help structure those scenarios; it does not guarantee employment or income. First get the workplace answers above, then decide whether a personal comparison would add anything.

Sources: FY 2026 ICD-10-CM Official Guidelines for Coding and Reporting; Science & Tech Spotlight: AI for Medical Notes and Coding; Medical Records Specialists: Occupational Outlook Handbook; Employment Projections Methods Overview; Occupational Separations

Sources and notes

  1. FY 2026 ICD-10-CM Official Guidelines for Coding and Reporting

    The official FY2026 guidelines apply October 1, 2025 through September 30, 2026. They state that the entire record should be reviewed to determine the encounter reason and conditions treated, stress consistent and complete documentation and provider-coder collaboration, and instruct users to follow the classification conventions and code-specific rules. They do not evaluate AI products.

  2. Science & Tech Spotlight: AI for Medical Notes and Coding

    GAO’s technology spotlight describes AI medical coding tools and potential uses and risks. It reports few independent accuracy studies and difficulty conducting independent assessments; inaccuracies may cause patient harm or over- or under-reimbursement. It is an evidence and policy overview, not a controlled comparison of vendors or a universal error-rate estimate.

  3. Extracting International Classification of Diseases Codes from Clinical Documentation Using Large Language Models

    The PubMed abstract reports that, among 50 inpatient notes, a human coder extracted 165 unique ICD-10-CM codes while six LLMs extracted 221 to 658; agreement was minimal to none overall. The counts do not by themselves identify which codes were erroneous, and the study does not establish performance in every current system or workflow.

  4. Can GPT-3.5 generate and code discharge summaries?

    The accessible University of Edinburgh repository copy of the 2024 JAMIA study reports generation of 9,606 synthetic discharge summaries and a separate direct-coding evaluation of GPT-3.5 on real MIMIC-IV notes. It reports F1 scores of 14.76% (leaf-only) and 17.13% (set-based), and concludes GPT-3.5 alone in the tested prompt setting was unsuitable for ICD-10 coding. These are study-specific results, not a current product error rate.

  5. ICD-10 | CMS

    CMS lists FY2027 ICD-10-CM files for encounters and discharges from October 1, 2026 through September 30, 2027, and separate FY2026 files for periods before and after the April 1, 2026 update. This supports checking the applicable code-set version against the encounter or discharge date, not claims about any vendor’s update schedule.

  6. Medical Records Specialists: Occupational Outlook Handbook

    BLS projects 8% employment growth and about 14,000 annual openings for U.S. medical records specialists from 2025 to 2035, and says many openings are expected from replacement needs. This is an occupation-wide national projection; it does not establish local vacancies, demand for particular adjacent duties, or AI effects on a specific job.

  7. Employment Projections Methods Overview

    BLS describes its employment projection process as six interrelated steps and explains that occupational employment projections are developed in the National Employment Matrix. It says occupational job openings include growth and separations; this overview describes method, not local vacancies or an AI effect for a particular role.

  8. Occupational Separations

    BLS says projected occupational openings combine projected employment change with estimated separations from workers transferring to other occupations or leaving the labor force; workers changing jobs within the same occupation are not counted. Its separations method uses historical patterns to project future separations.

Apply this to your own work

See the whole job market at once.

Explore which occupations AI may reshape, then turn the signal into a practical response.

Explore the job map