ClinicOps / Briefings / Methodology
Methodology · Published Jun 2, 2026
How We Collect Benchmark Data (Methodology and Anonymization)
A benchmark is only as trustworthy as how it was made. So here is exactly how our benchmark data is collected, why no patient information ever touches it, how it is anonymized, and what we will and will not publish. We show our work, because numbers you cannot check are numbers you should not trust.
Our benchmark methodology rests on three commitments: the data is operational, never patient information, so no PHI ever enters it; it is always reported in aggregate, with no practice named or made identifiable and no segment too small to protect anonymity; and every figure is labeled by source. The first report compiles published industry benchmarks; our own aggregated audit data is added over time, clearly marked as such.
Key takeaways
- The data is operational (denial rates, AR, prior auth hours), never protected health information.
- Chart-numbers-only applies here as everywhere, so no patient data enters the dataset at any point.
- Everything is reported in aggregate; no practice is ever named or made identifiable.
- No segment is published from a sample too small to protect anonymity.
- Every figure is labeled by source: industry benchmarks now, our own aggregated audit data added over time.
Benchmark reports are easy to publish and hard to trust, because most of them never show how the numbers were made. A figure with no visible origin is just an assertion. So before you use our benchmarks, in the benchmark report, here is exactly how they are produced, including the parts that protect the practices behind the data. If we cannot explain where a number came from, we do not publish it.
Why methodology matters
A benchmark is a claim about what is normal, and a claim is only as good as its basis, so the methodology is not a technical footnote; it is the thing that makes the number worth anything. When a report tells you the typical denial rate without saying where that figure came from, how it was gathered, from whom, over what period, you have no way to judge whether it applies to you or whether it is trustworthy at all, so you are taking it on faith. We would rather you did not have to. Showing the methodology does two things: it lets you evaluate the numbers rather than accept them, and it disciplines us, because committing to publish only figures whose origin we can explain rules out the vague, unsourced statistics that clutter the field. This is the same transparency principle we apply to prices and to the offshore reality of how we work, extended to data: show the real thing, including the limits, rather than a polished number with a hidden basis. Methodology transparency is how a benchmark earns trust instead of just asserting it.
What we collect, and what we do not
The most important line in our methodology is the one about what never enters the data. We collect operational metrics, not patient information. The figures are practice-level: denial rates, days in accounts receivable, prior authorization hours, no-show rates, clean claim rates, and similar measures of how a practice runs, the same numbers on the KPI dashboard. What we do not collect is any protected health information, no patient names, no diagnoses, no chart contents, nothing that identifies or describes a patient. This is not a special data-handling regime invented for benchmarking; it is the exact chart-numbers-only discipline we apply to every system we build, described in the HIPAA-safe project management guide. Operational metrics are aggregate properties of a practice's workflow, and they simply do not require patient data to compute, so patient data never enters the pipeline. The result is a dataset that can say meaningful things about how practices operate while containing nothing that could ever expose a patient, which is exactly how it should be.
How anonymization works
Protecting patients is the first commitment; protecting the practices behind the data is the second, and it is built in three ways. First, aggregate-only reporting: every published benchmark is pooled across many practices, so no figure represents or reveals any single practice. Second, no small-sample segments: we do not report a benchmark for a slice, a specialty, a region, a size band, unless the sample is large enough that no individual practice could be inferred from it, because a segment of two or three practices is not a benchmark, it is a privacy risk. Third, no practice is ever named or made identifiable, in any report, under any circumstances. The combined effect is that our benchmarks tell you what is normal across practices while making it impossible to trace any number back to a specific one. A practice that contributes data sees its own numbers, privately, and contributes only to pooled figures that protect it. Anonymization here is not a promise to be careful; it is a structural property of only ever publishing aggregates above a minimum sample size.
The free Leak Audit gives you your practice's metrics against benchmark, kept private to you.
Start with a free Leak AuditSourcing and labeling
Honesty about where each number comes from is the third pillar, and it matters especially right now, because our own dataset is still growing. So we label plainly. The first benchmark report compiles published industry benchmarks from established sources on denials, accounts receivable, prior authorization, staffing, and no-shows, and it says so; it does not dress up industry figures as proprietary data we gathered. As our own aggregated audit data grows, from the practice audits and engagements we run, with consent, future editions will add those findings alongside the industry figures, each clearly marked by source, so you always know whether a number is industry data or ours. We will not blend the two into an unlabeled composite, and we will not claim a proprietary dataset larger or more mature than it is. When our data is thin on a metric, we will say so; when it is robust, we will show it. This labeling discipline is what lets the report grow honestly over time, from a compilation of good industry data today into a genuinely proprietary benchmark as the dataset matures, without ever overstating what we have at any point along the way.
The limits, stated plainly
Trustworthy data means being as clear about what the numbers cannot do as about what they can, so here are the honest limits. Our own dataset is young. Today's report leans on compiled industry benchmarks precisely because our proprietary data is still accumulating, and we would rather say that than imply a mature dataset we do not yet have. Industry benchmarks are themselves imperfect, drawn from varied sources with different definitions and methods, so they are directional, not exact, which is why we present ranges and targets rather than false-precision single figures. Benchmarks are averages, and you are not an average. Your specialty, payer mix, and circumstances shift what is realistic, so a benchmark is a reference point for finding your gaps, not a verdict on your practice, a caveat we make explicit in the report itself. And a benchmark cannot tell you why. It can show that your denial rate sits above the norm; it cannot tell you the cause, which is what an actual look at your operations is for. Stating these limits is not hedging; it is the difference between data that helps you think and data that pretends to more authority than it has. We would rather give you honest numbers with clear limits than confident numbers you cannot rely on.
The standards we hold
Underneath the specifics sit a few standards we hold on every figure we publish. Every number is sourced and dated, so you can see where it came from and when, the same rule we apply to every statistic in these guides. We do not cherry-pick, selecting flattering figures or dropping inconvenient ones to tell a cleaner story. We report honest sample sizes and limits, rather than implying more certainty than the data supports. And we do not manufacture precision, publishing a suspiciously exact figure to seem authoritative when the honest answer is a range. These are not onerous standards; they are just what taking data seriously requires, and they are the same commitments, real numbers, no fabrication, no manufactured authority, that govern everything else we publish, from case studies to prices. The point of laying all this out is simple: a benchmark you cannot check is a benchmark you should not trust, so we show you how ours are made, what protects the patients and practices behind them, and where they come from. Then the numbers in the report are yours to judge, which is exactly how it should be.
Where to go next
- The State of Independent Practice Operations: Benchmark Report #1 live
Benchmark report for independent practices: denial rate, days in AR, prior auth hours,.
- Practice Manager KPIs: The 12 Numbers to Track Weekly (Dashboard Template) live
The 12 KPIs a practice manager should track weekly: AR, clean claim rate, denials, no-shows,.
- HIPAA, ClickUp, Monday and Asana: The PHI-Safe Setup (Chart Numbers Only) live
Run ClickUp, Monday, or Asana in a HIPAA-safe way: keep PHI out with the chart-numbers-only.
Find the leak before you fix it
Two ways to start, both free.
Run the free Rescue Kit and its tools yourself, or book a 20-minute Leak Audit where we put a real number on what this is costing, using your own volume. A diagnosis, not a pitch.
Frequently asked questions
How is the benchmark data collected?
This first report compiles published benchmarks from established industry sources. As our own dataset grows, it will draw on operational metrics gathered through practice audits and engagements, with consent, always aggregated and never tied to an identifiable practice. Every figure is labeled by source so you know whether it is industry data or ours.
Does the benchmark data include patient information?
No. The data is operational, denial rates, days in AR, prior authorization hours, no-show rates, and similar practice-level metrics, never protected health information. We apply the same chart-numbers-only discipline to our data that we apply to every tool we build, so no patient data enters the dataset at any point.
How is the data anonymized?
By design and by policy: we collect practice-level operational metrics, not patient data, we report only in aggregate, and we do not report any segment small enough to risk identifying an individual practice. No practice is ever named or made identifiable in a benchmark, and no figure is published from a sample too small to protect anonymity.
What sources does the first benchmark report use?
Published industry benchmarks from established sources on denials, accounts receivable, prior authorization, staffing, and no-shows. It is labeled as such. Future editions will add our own aggregated audit findings alongside the industry figures, each clearly marked by source, rather than blending them or presenting industry data as proprietary.
Why does methodology transparency matter?
Because a benchmark is only as trustworthy as how it was produced. Showing the sources, the anonymization, and the limits lets you judge the numbers rather than take them on faith, and it keeps us honest, since a figure we cannot explain the origin of is a figure we should not publish. Transparency is the point, not a footnote.
Will you ever share individual practice data?
No. Individual practice data is never shared, published, or made identifiable, full stop. Benchmarks are always aggregate. The only numbers that ever appear are pooled across enough practices to protect anonymity, and any practice's own data stays private to that practice.