Blog

Audit season: what payers actually pull

A post-payment review is not an accusation, it is a sampling exercise with money attached. The payer already paid you, because paying is cheaper than reviewing; now analytics have flagged something, a small set of charts is being read by a stranger, and whatever error rate that sample produces will be applied to everything that looks like it. Understanding the mechanism is most of the defence.

Payers review after they pay, not before, and they review by sample. A claim is auto-adjudicated in seconds against format and coverage rules; nobody reads the chart. Months or years later, analytics flag a provider whose pattern differs from peers, a records request goes out for a set of dates of service, a reviewer reads those charts against medical-necessity criteria, and the error rate found in the sample is projected across the comparable claim universe. That projection, not the sampled claims, is where the money is.

How does post-payment review actually work?

Four stages, and they are worth separating because different things are true at each one.

Adjudication. The claim arrives, passes format and eligibility edits, matches a covered code and a covered diagnosis, and pays. This step contains no clinical judgement. A paid claim means nothing about whether the documentation supports it, which is the thing practices most consistently misread as approval.

Analytics. The payer compares you against peers on distributions: code mix, units per client per year, session length patterns, diagnosis concentration, telehealth share. Outliers get flagged. Being an outlier is not evidence of anything; a trauma-focused practice doing exposure work genuinely has a different code distribution than a general practice. It is simply what triggers a look.

Records request. A letter asking for charts on a specific list of dates, typically within 30 days. This is the stage where your existing documentation becomes fixed: whatever is in the chart when that letter arrives is what you have.

Determination and demand. Findings per claim, an error rate, and, if the programme uses it, extrapolation to the wider universe. Then an appeal window, usually short, running from the determination date.

Federal programme integrity work is documented publicly at oig.hhs.gov and the Medicare-side process at cms.gov. Commercial payers run structurally similar programmes under contract terms rather than under regulation, which matters because your rights in a commercial review come from your provider agreement, not from statute.

What a records request looks like

The letter is dull and specific. It names the payer or a delegated review vendor, gives a reference number, lists dates of service with member IDs and claim numbers, states a return deadline, and specifies document types. That last part is the part to read closely, because it usually asks for more than notes.

  • Progress notes for each listed date of service
  • The treatment plan in force on each of those dates
  • The intake or diagnostic assessment establishing the diagnosis
  • Signed consent to treat, and consent for telehealth where applicable
  • Any authorisation on file
  • Assessment instrument scores where they support severity
  • Occasionally: the appointment schedule, to corroborate that sessions occurred as billed

The rule that matters most. Do not edit existing notes. A payment dispute is expensive; an altered record is a different category of problem entirely, and audit trails in modern systems record edits with timestamps. If something genuinely must be added, add it as a clearly labelled late entry carrying the date it was actually written. Never overwrite.

What the reviewer checks first

A reviewer with a stack of charts and a quota does not read your notes as clinical literature. They work a checklist, and the first item is the one most notes fail: does this session connect to an active, measurable goal in a treatment plan that was in force on that date.

That link is the golden thread: diagnosis to treatment plan goal to session note to billed code, as one continuous chain. Break any link and the chart stops demonstrating treatment and starts demonstrating a conversation. Payers do not reimburse conversations at psychotherapy rates.

The common failures are structural rather than clinical. A treatment plan signed after the sessions billed against it. A plan whose review date passed eleven months ago. Goals written as aspirations ("reduce anxiety", "improve coping"), which cannot be evaluated as met or unmet. A note that describes a genuinely good session, in detail, without ever naming which goal it advanced. See treatment plans and progress notes for what each document has to carry.

Time documentation, and why it fails

Psychotherapy codes are timed, and time is the easiest finding a reviewer can make because it requires no clinical judgement at all. Either the minutes are documented or they are not.

Three patterns produce findings. The first is simply no time recorded: the note describes an hour of work and never says how long it took. The second is a default: every note for eighteen months says exactly "60 minutes", which reads as a template value rather than a measurement, because human sessions do not run to identical durations for a year and a half. The third is a mismatch between the recorded time and the code billed, which is arithmetic and therefore not arguable.

Count face-to-face time only. Note writing, records review and care coordination are real work, some of it separately billable, and none of it converts a shorter session into a longer code. CPT 90837 lays out the bands; the AMA maintains the code set itself at ama-assn.org.

The 90837 pattern problem

The individual psychotherapy codes split at 53 minutes: 38 to 52 minutes is 90834, 53 minutes and above is 90837, and 90837 pays meaningfully more. One minute of documented time changes the reimbursement, which is exactly why payers built analytics around this pair.

Here is the shape of the problem. A practice billing mostly 90834 with 90837 where clinically indicated produces a distribution that looks like clinical variation. A practice billing 90837 on 95% of sessions produces a distribution that looks like a billing decision, and payer analytics are specifically built to find billing decisions. Neither pattern proves anything, and an EMDR-heavy trauma practice may legitimately run near-universal 90837, but the second one gets read, and once it is read the charts need to independently justify the intensity.

The two pieces of advice you will find elsewhere are both wrong. "Avoid 90837, it triggers audits" tells clinicians to under-bill for work they did. "Bill what you did and ignore the noise" is correct right up until the records request, at which point the note has to carry the weight alone. The narrow correct position: bill 90837 whenever the session genuinely ran 53 minutes or more, record the actual minutes rather than a default, and make sure the note names the goal and the intervention that justified that length.

How to respond in 30 days

Thirty days is enough to assemble a defensible package and not enough to change what is in the charts. Work in this order.

  1. Days 1–2: scope it. Log the deadline, the reference number, the exact date list, and the document types requested. Set your internal target at deadline minus seven.
  2. Days 3–10: assemble. For every listed date, pull the note, the plan in force on that date, the intake, the consent, the authorisation and any instrument scores. Assemble by date of service, not by document type; reviewers read by date.
  3. Days 8–14: reconcile. For each date, check billed code against documented minutes, and check that the note names an active goal. Write down every discrepancy you find. You are not fixing them; you are learning what the reviewer will find, which determines whether you need help.
  4. Days 12–18: decide on representation. If the letter mentions extrapolation, a statistical sample, or a projected overpayment, get a healthcare attorney or a certified coder involved now rather than after the determination. The cost asymmetry is not close.
  5. Days 18–23: write the cover letter. One page. Reference number, dates enclosed, an index of what is supplied per date, and a short factual statement of your documentation practice. No arguments about clinical judgement; that is for the appeal, if there is one.
  6. Days 23–25: send with proof. Portal upload with a saved receipt, or tracked mail. Keep an exact copy of the package as sent, in the order sent.
  7. Day 25 onward: calendar the appeal. Note when a determination is expected and what the appeal deadline would be. Appeal windows are short and start at determination, not at the point you opened the envelope.

A printable version of that sequence is on the templates page, alongside note structures and a treatment plan skeleton.

Why extrapolation is the whole game

The arithmetic is worth doing once because it changes how the request feels. Suppose the sample is 30 claims and 12 are found deficient: a 40% error rate. Suppose the comparable claim universe over the lookback period is 920 claims at an average allowed amount of $110, a figure chosen for the illustration, since real contracted rates are confidential. The universe is $101,200. Applying 40% produces a demand of about $40,480, from a sample whose face value was 30 × $110 = $3,300.

Twelve charts turned into forty thousand dollars. That ratio is the reason routine documentation habits matter more than any single case, and the reason step three above exists: knowing your own error rate before the reviewer calculates it tells you exactly how much this is worth spending on. Clawbacks and recoupment covers the appeal path, the offset mechanism payers use to collect while an appeal is pending, and where sampling methodology itself is challengeable.

What to change before the letter arrives

Everything above is a consequence of three habits, none of which cost anything to adopt and all of which are expensive to retrofit across two years of charts. Record actual session minutes rather than a template default. Name the treatment plan goal in every note, in the assessment section, in a sentence. Review and re-sign treatment plans on schedule, with the client's participation documented.

Those three make a chart survivable. They are also, not coincidentally, the parts of the record that are hardest to keep aligned when the plan lives in one document, the note in another and the code in a third, which is a structural problem rather than a discipline problem. Weft puts the plan goal on the note as a field and the code beside the documented time, so the thread is a consequence of the workflow rather than something to remember at 7pm.

Questions

Common questions

How far back can a payer audit claims?

It depends on the payer and the programme. Commercial lookback periods come from your provider agreement and commonly run one to three years; government programmes have their own statutory windows, and allegations of fraud extend them considerably. Check the recoupment and audit clauses in your contract, because that is the binding answer for commercial plans.

Does a paid claim mean the documentation was accepted?

No. Adjudication is automated and checks format, eligibility, code coverage and diagnosis coverage. No human reads the chart at payment. Post-payment review exists precisely because payment happens first, so a long history of clean payments tells you nothing about whether your notes would survive a reviewer.

Can I add a missing note after a records request arrives?

You can add a late entry, clearly labelled as such and carrying the date it was actually written. You cannot edit or backdate an existing note. Systems record edit timestamps, and an altered record converts a payment dispute into a much more serious allegation. Send what exists and disclose the gap.

Is billing 90837 frequently going to trigger an audit?

Frequency alone does not prove anything, but distribution is what analytics compare. Near-universal 90837 draws a look, and once the charts are read they have to justify the length independently. Bill it whenever the session genuinely ran 53 minutes or more, record the real minutes, and name the goal that justified the intensity.

What is extrapolation in a payer audit?

The payer reviews a sample of claims, calculates an error rate, and projects that rate across the full universe of comparable claims in the lookback period. A 40% error rate found in 30 charts can produce a five-figure demand from a sample worth a few thousand dollars. If extrapolation is mentioned, get professional representation.