MelloJam Listen now

10 Patient Satisfaction Statistics Every Clinic Manager Should Track

Patient satisfaction is not one universal percentage. It can describe a patient's overall evaluation, a report of what happened during care, a rating of a provider, or an answer to one specific question about access. Those measures are useful only when the question wording, response scale, denominator, survey version, and reporting level stay visible.

We read the original AHRQ CG-CAHPS guidance, its supplemental access items, the CAHPS data tools and calculation rules, and an outpatient scheduling study. We extracted the exact measure groups, wait-time questions, top-box definitions, comparison boundaries, and source limitations so clinic operators can build a patient experience reference that is specific enough to use.

Back to Clinic statistics

1. Patient experience, satisfaction, quality, and outcomes are different measures

Patient experience asks what the patient reports happened, such as whether the provider explained things clearly or whether the office kept the patient informed about a delay. Patient satisfaction is an evaluation of that experience, often expressed as an overall rating or a judgment about a particular part of the visit. A provider rating is a global score about the named provider. These measures overlap, but they do not mean the same thing.

Clinical quality and health outcomes are different again. A high top-box score for communication does not prove that care was clinically effective, and a low rating does not by itself identify a clinical error. The AHRQ calculation guidance treats patient experience scores as reports and ratings, not as a substitute for clinical outcomes. A defensible article should keep those categories separate from the first paragraph.

2. The five core CG-CAHPS measures cover access, communication, coordination, staff, and provider rating

AHRQ's Clinician and Group Survey is designed for patient reports about providers and staff in primary and specialty care. Its core measure groups give a clinic a practical starting dictionary:

The exact item set depends on the survey version and population. The table below keeps the core measure names beside the response scales and the top-box rule that makes the score reproducible.

Domain Exact item or composite Response scale Top-box definition Denominator Reporting level Source
Access Getting Timely Appointments, Care, and Information Usually four-point frequency items: Never, Sometimes, Usually, Always Average the item-level Always proportions for the composite Valid responses for each item; missing responses excluded from each proportion Respondent or database level; practice site and group reporting when available AHRQ CG-CAHPS guidance
Communication How Well Providers Communicate With Patients Usually four-point frequency items Always for each item; composite top box is the average of item proportions Valid item responses Respondent or database level; practice site and group reporting when available AHRQ CG-CAHPS guidance
Care coordination Providers' Use of Information to Coordinate Patient Care Usually four-point frequency items Always for each item; composite top box is the average of item proportions Valid item responses Respondent or database level; practice site and group reporting when available AHRQ CG-CAHPS guidance
Office staff Helpful, Courteous, and Respectful Office Staff Usually four-point frequency items Always for each item; composite top box is the average of item proportions Valid item responses Respondent or database level; practice site and group reporting when available AHRQ CG-CAHPS guidance
Provider rating Patients' Rating of the Provider Global rating from 0 to 10 Ratings of 9 or 10 Valid global ratings Respondent or database level; site and group comparisons require their own reporting rules AHRQ results calculation guide
Wait timeliness AC9: saw this provider within 15 minutes of the appointment time Never, Sometimes, Usually, Always Always Valid AC9 responses Practice site, provider, or group, if the sampling design supports that unit AHRQ supplemental access item AC9
Wait communication AC10: was kept informed about how long the patient would need to wait after checking in Never, Sometimes, Usually, Always Always Valid AC10 responses Practice site, provider, or group, if the sampling design supports that unit AHRQ supplemental access item AC10

Treat the table as a measure dictionary, not a promise that every version uses identical wording. A clinic should record the survey version and the full item wording in its data dictionary before publishing a number.

3. Two exact access questions separate being on time from being informed about a delay

AHRQ's adult 3.0 supplemental access set contains two questions that are especially useful for clinic operations. AC9 asks: "Wait time includes time spent in the waiting room and exam room. In the last 6 months, how often did you see this provider within 15 minutes of your appointment time?" AC10 asks: "In the last 6 months, after you checked in for your appointment at this provider's office, how often were you kept informed about how long you would need to wait for your appointment to start?"

Both questions use the response options Never, Sometimes, Usually, and Always. The top-box result is the percentage choosing Always among valid responses. AC9 measures whether the patient experienced a timely start under the survey's stated definition. AC10 measures communication about the wait. A clinic can perform poorly on one and well on the other, so combining them into one satisfaction number destroys a useful operational distinction.

The same access set also offers concrete time-to-appointment categories. AC1 asks about the usual wait for care needed right away, with choices from Same day through More than 7 days. AC2 asks about routine care, with choices from Same day through More than 30 days. Use these as separate access measures and keep them separate from the in-clinic wait questions. Supplemental items should be used only when the sample design is likely to produce enough responses for analysis and reporting.

4. Top-box means the most positive valid response, not the average response

The basic item formula is:

item top-box score = respondents choosing the most positive response / valid responses to the item x 100

AHRQ's calculation guide excludes missing responses from top-box and proportional-score percentages. Its response crosswalk is:

Response scale Lower proportion Middle proportion Top-box response
Dichotomous No - Yes
Global rating 0 to 6 7 to 8 9 to 10
Three-point No Yes, somewhat Yes, definitely
Four-point Never or Sometimes Usually Always

For example, if 4 of 10 valid respondents choose Always, the item top-box score is 40%. It is not the mean of the four response categories, and it is not 40% of all people who were invited if some responses are missing. Report the valid response count beside the percentage so readers can see whether a score rests on 10 responses or 1,000.

5. A composite score averages item proportions, so it is not simply a count of satisfied patients

For a composite measure, AHRQ calculates the proportion in each response category for every item and then averages those proportions across the items. For the top box, the simplified formula is:

composite top-box score = sum of item top-box percentages / number of items

The AHRQ calculation guide gives a concrete example for Getting Timely Appointments, Care, and Information. If its three item top-box scores are 68%, 72%, and 61%, the composite top-box score is 67%, calculated as (68 + 72 + 61) / 3.

That 67% does not necessarily mean that 67% of patients answered Always to every question. It is the average of three item-level proportions. This distinction matters when a clinic explains why a composite moved: the change may come from one item, several items, or a change in which items had valid responses.

6. Survey version, reference period, and reporting level can change the meaning of a comparison

AHRQ lists three relevant survey versions. The Clinician and Group Visit Survey 4.0 beta asks about the most recent synchronous visit, whether it happened in person, by phone, or by video. Versions 3.0 and 3.1 ask about care over the last 6 months. The 3.0 and 3.1 sampling frame includes patients with at least one scheduled or unscheduled visit in that period. A most-recent-visit score and a six-month score should not be presented as though they used the same reference period.

The unit of analysis matters just as much. AHRQ says the survey can be administered at the provider, practice-site, or group level. Sample at the practice-site level when the question is site performance, and at the provider level when the question is individual clinician performance. A group-level score that combines sites can hide the variation a manager needs to find.

The CAHPS data tool reports aggregated Clinician and Group results for survey years 2018 through 2019, and AHRQ says new submissions to that database were suspended in 2021 because of declining participation. Treat those results as a historical comparison resource, not as a live 2026 national benchmark. Every comparison should state the survey version, reference period, survey year, and reporting level.

7. Small samples can suppress a result, so a blank value is not zero

The AHRQ calculation guide defines a complete record as responses for at least 50% of key items plus a response for one or more core composite or rating items. A partial complete record has a response for one or more core composite or rating items. This is a data-inclusion rule, not a satisfaction judgment.

The same guide gives suppression rules that should be visible in any clinic dashboard:

For a local survey, define response rate = completed survey responses / eligible sampled patients x 100 and publish the eligibility rule beside it. A low response rate can make a score less representative even when the calculation is mathematically correct. AHRQ also warns that handing a survey to a patient during or immediately after an office visit may create more positive ratings and lower response rates than mail or telephone administration, so collection mode belongs in the methodology notes.

8. Case-mix adjustment affects significance tests, but not the public top-box percentage

AHRQ's calculation guide says group and practice-site mean scores are case-mix adjusted before statistical difference testing. The listed adjustment characteristics are age, education, and self-reported general and mental health status. The top-box scores and percentiles shown in the public Online Reporting System are not case-mix adjusted.

This creates two different comparison questions. A raw top-box score tells you the proportion of valid responses in the most positive category. A significance result asks whether an adjusted mean differs from a comparison mean after accounting for selected respondent characteristics. A site can therefore have a top-box result that does not line up intuitively with a significance flag. Do not treat the flag as a more precise version of the top-box number.

A percentile is also relative, not a percentage of patients satisfied. Under the AHRQ definition, a 75th percentile is the top-box score at or below which 75% of practice-site top-box scores fall. It does not mean that 75% of the site's patients chose the top response.

9. One outpatient study linked waiting, service time, and clinic environment to satisfaction in a bounded setting

The outpatient study provides a useful set of reported figures, but it is not a national patient-experience survey. It covered 10 outpatient clinics with different specialties in Mashhad, Iran. The clinics had no web-based appointment system, operated only during the evening shift, and contributed data from December 2016 through March 2017. A total of 319 patients completed the questionnaire.

Reported measure Result Study denominator or scale How to use it
Overall patient satisfaction 6.73 plus or minus 0.16 Visual analog scale reported on a 0 to 10 range A bounded example of an overall score, not a universal target
Clinic environment satisfaction 8.30 plus or minus 0.12 Visual analog scale reported on a 0 to 10 range A signal that environment was part of the experience in this study
Average waiting time 64.2 plus or minus 3.45 minutes 319 patients across 10 clinics; reported range 0 to 302 minutes A source-specific flow benchmark, not a national average
Average service time 9.85 plus or minus 0.37 minutes 319 patients across 10 clinics; reported range 1 to 35 minutes Physician time with the patient under the study's definition
Patients satisfied with wait length 57.7% 184 of 319 patients A categorical satisfaction result from this questionnaire
Patients satisfied with service duration 79.9% 255 of 319 patients A categorical satisfaction result from this questionnaire

The authors found statistically significant associations between overall satisfaction and satisfaction with the clinic environment, waiting time, and service time. The study also reports that 139 patients, or 43.7%, expressed overall satisfaction in its reporting. Because the questionnaire used both visual analog and categorical measures, those figures should not be merged into one invented score.

The operational takeaway is a question to test locally: are patients reacting to the length of the wait, the service itself, the environment, or whether staff explain the delay? The study supports asking that question. It does not prove that every clinic should target 64 minutes, 9.85 minutes, or any particular satisfaction percentage.

10. The most useful clinic dashboard pairs every patient score with its question, denominator, and operational signal

A clinic can make its satisfaction reporting linkable and actionable by keeping the survey result beside the operational measure it might explain:

Dashboard measure Exact definition Required denominator or context Operational companion
Timely appointment top box Percentage selecting the most positive valid response for the relevant access item or composite Valid item responses, survey version, and reference period Days to urgent and routine appointment
Seen within 15 minutes top box AC9 respondents selecting Always Valid AC9 responses; state that wait includes waiting and exam room time Timestamped scheduled start and provider start
Wait communication top box AC10 respondents selecting Always Valid AC10 responses; state that the item asks about the period after check-in Wait-notification process and actual wait interval
Provider communication top box Valid responses selecting the top response on the communication items Item-level valid n and composite calculation Complaint themes, clarification requests, or follow-up contacts
Office staff top box Valid responses selecting the top response on staff items Item-level valid n and survey mode Check-in, checkout, and call-handling workflow
Provider rating top box Valid global ratings of 9 or 10 Valid 0 to 10 ratings; provider, site, or group level Provider-level visit mix and service-time distribution
Response rate Completed responses divided by eligible sampled patients Eligibility definition, collection mode, and partial-completion rule Outreach and survey distribution process
Percentile comparison Position of a site score within a defined comparison distribution Survey year, version, comparison population, and site count Internal trend over time, not a clinical outcome

Start with the exact question and the valid response count, then add the interpretation. A survey score without its response scale is hard to compare. A top-box score without its denominator is hard to trust. A satisfaction score without an operational companion is hard to improve. Keeping those three layers together turns patient experience data into a durable reference page rather than a collection of vague claims.

Sources