Written one- and two-star Google reviews of at least five words, up to 60 per business, from high-review-count veterinary practices in eight US metros. A language model, with no human-coded check, classified the main reason in 9,149 of them: 70.6 percent how the business is run, 28.1 percent the veterinary medicine itself, 1.3 percent unclear.
The study covers 9,154 written one- and two-star Google reviews of at least five words, drawn from the up-to-60 lowest-rated reviews retrieved per business, from 296 high-review-count veterinary practices in Phoenix, Dallas, Atlanta, Charlotte, Columbus, Denver, Tampa and Kansas City. In each metro, up to 200 listings were screened and, after filters including a cap of two locations per brand, the 40 businesses with the most Google reviews were sampled. Exclusions then left 296 businesses, of which 296 have reviews in the corpus. Review dates run from 2008-08-11 to 2026-09-30. A separate multi-label model pass assigned at least one business-process label to 93.1 percent of the 9,154 reviews.
The model assigned at least one core-service label to 39.6 percent of the 9,154 reviews, and the model assigned the clinical quality label to 35.8 percent. Every figure below states its denominator, and none of them is a share of clients or patients.
The model-classified main reason
The main-reason pass was run by a language model, Claude Haiku 4.5, which chose between how the business is run, the veterinary medicine itself, and unclear. It classified 9,149 of the 9,154 reviews.
| Model-classified main reason | Share of the 9,149 reviews labelled by the main-reason pass |
|---|---|
| How the business is run | 70.6% |
| The veterinary medicine itself | 28.1% |
| Unclear | 1.3% |
A second model, Claude Sonnet 4.5, classified 200 of the same reviews. On the binary question of how the business is run against every other main reason, the two models agreed on 92.5 percent of the 200, with a Cohen’s kappa of 0.83 (95 percent confidence interval 0.74 to 0.90). It is not validation against human coding, and no sample was checked against human coding.
Process and the medicine, counted label by label
A separate pass assigned theme labels, and a review can carry more than one. Some describe the business process, such as pricing, scheduling or communication. Two describe the medicine and the handling of the animal: clinical quality and animal handling. A single review can carry both kinds.
| What the labels cover | Share of the 9,154 reviews |
|---|---|
| A business-process label only | 57.6% |
| Both a process label and a core-service label | 35.6% |
| A core-service label only | 4.0% |
| Neither a process nor a core label | 2.9% |
The multi-label pass and the main-reason pass answer different questions. One asks what a review mentions, the other asks what the model classified as the main reason, and the two are not reconciled review by review.
The model-assigned themes
The table lists every label in the taxonomy. These are themes a model assigned to what reviewers wrote, not verified events. A review can carry several labels. The shares overlap and are not summed.
| Model-assigned theme | Share of the 9,154 reviews |
|---|---|
| Staff conduct | 45.5% |
| Communication failure | 40.9% |
| Pricing transparency | 38.3% |
| Clinical quality | 35.8% |
| Scheduling and access | 22.6% |
| Upsell pressure | 16.5% |
| Refused service | 14.0% |
| Other | 6.9% |
| Animal handling | 6.4% |
| Facility conditions | 2.5% |
Staff conduct
The model assigned the staff conduct label to the largest share, 45.5 percent of the 9,154 reviews. It is also one of the two catch-all process labels in the taxonomy. Both quoted excerpts describe front-desk or reception conduct.
“Felt like the front desk person was rude and not helpful and was trying to avoid us to talk to anyone.”
“he was rudely told he was late and that the appointment was cancelled. When he told him the appointment was at 2:30, and the discrepancy was discovered, no apology was given.”
Communication failure
The model assigned the communication failure label to 40.9 percent of the 9,154 reviews. It is the second catch-all process label. The quoted excerpts describe a requested callback that did not come and a long period with virtually no communication before the reviewer sought an update.
“I called twice and asked for a call back. They couldn’t be bothered.”
“we were left with virtually no communication for over an hour. I had to go outside and flag down random staff just to get an update”
Pricing transparency
The model assigned the pricing transparency label to 38.3 percent of the 9,154 reviews. Both quoted excerpts compare an earlier quoted price with a later, higher figure.
“After original quote of $582 plus extractions, I was expecting to pay between $800-$900. When I got there the estimate was $1500 to $2000.”
“I was quoted $169 for a few services when I got to the front counter to pay, it went up to $214.”
The amounts are the reviewers’ own figures. The study did not check any estimate or invoice, and makes no finding about whether any charge was justified.
Scheduling and access
The model assigned the scheduling and access label to 22.6 percent of the 9,154 reviews. The quoted excerpts describe no sick visits being available and a wait in the clinic.
“was told that they have no sick visits for the next 4 weeks”
“Waited over 35 minutes for my dogs to be taken back for an appointment”
Disclosure: this study was produced by Sales Roadmaps, which sells operations consulting services. About the operations consulting service
The number that keeps this honest
The 296 businesses sampled after exclusions hold 227,903 lifetime Google reviews between them. Of those, 15,220, or 6.7 percent, are one or two stars, and 86.7 percent are five stars. All 296 have at least one review in the corpus studied here.
Percentages in this study use the denominator stated with each figure. Main-reason shares are of the 9,149 reviews the main-reason pass labelled. Label shares are of all 9,154 reviews. The base rate and the five-star share are of all 227,903 lifetime reviews. None of them is a share of clients or patients.
Operational questions these reviews raise
These are questions for an owner reviewing their own operation. They are not findings. The study reports what reviewers wrote and how models labelled those reviews. Related reading on Sales Roadmaps covers building a price book, CRM setup and business process management.
Which updates are promised during a visit?
The model assigned the communication failure label to 40.9 percent of the 9,154 reviews. The quoted examples describe a requested callback and a long period with virtually no communication before the reviewer sought an update, and the study did not measure how often either situation occurs. An owner can ask which updates the practice promises during a visit or procedure, and who gives them.
How is a changed estimate shown before checkout?
The model assigned the pricing transparency label to 38.3 percent of the 9,154 reviews. The quoted examples describe a later figure higher than an earlier quote, a situation the study did not count separately. An owner can ask how a change to an estimate is shown before checkout.
What happens when no appointment is available?
The model assigned the scheduling and access label to 22.6 percent of the 9,154 reviews. The quoted examples describe no sick visits being available and a wait in the clinic. An owner can ask what message the practice uses when no appointment is available.
How the study was done
Google reviews were retrieved through the DataForSEO business data API, sorted by lowest rating, up to 60 per business. In each of the eight metros, up to 200 listings were screened. After filters, the 40 businesses with the most reviews per metro were sampled, with at most two locations per brand.
Twenty-four businesses were excluded. Six had the vendor primary category emergency veterinarian service. Eighteen were excluded after their names and listings were read: 9 specialty or emergency referral hospitals, 6 non-profit or low-cost spay and neuter clinics, 1 shelter, 1 in-home hospice and euthanasia service and 1 university teaching and referral hospital.
Those exclusions did not remove all emergency or specialist vocabulary. In 14.3 percent of the 9,154 reviews, 1,306 in total, the text matches a case-insensitive keyword pattern for emergency, specialist, referral, oncology, neurology or after-hours vocabulary. That is a keyword match, and what each of those reviews is about was not checked.
The corpus consists of the retrieved reviews rated one or two stars with at least five words, after exclusions and deduplication. Identical review texts posted under two listings were counted once, which removed 1 review. Candidate review-theme codes were proposed by a language model, Google Gemini 3.1 Flash-Lite, on a seeded sample of 200 reviews, and the research team consolidated them into a fixed label set.
Labels were then assigned by the same model. A second model, DeepSeek V3.2, relabelled 200 reviews and agreed on process against not process in 97.0 percent of them, with a kappa of 0.796. Each label was asked to carry a verbatim span from the review, and 94.9 percent of labels carry one. The rest are counted without a span.
As a label-based sensitivity test, which is not a main-reason result, the two catch-all process labels, staff conduct and communication failure, were ignored. On that basis, 70.3 percent of the 9,154 reviews still carry at least one process label. Reweighted by each business’s lifetime one- and two-star count, the any-label process share is 93.8 percent. With equal weight per business it is 92.4 percent.
The stored review data carry a business key, the star rating, the date, the text and an owner-reply flag only. No reviewer names were stored. No reviewer or reviewed business is named in this article.
Limitations
Google reviews are self-selected. The sample covers high-review-count businesses in eight metros: after filters including a cap of two locations per brand, the 40 businesses per metro with the most Google reviews among up to 200 listings screened, before exclusions. Review dates run from 2008-08-11 to 2026-09-30, and the sample is sorted by rating, not by date, so the study makes no claim about trends over time.
All labels and main reasons were assigned by language models, and no sample was checked against human coding.
Contact: Get in touch with Sales Roadmaps
Frequently Asked Questions
In this study, which model-assigned themes are most common?
Staff conduct, which the model assigned to 45.5 percent of the 9,154 reviews studied. Communication failure follows at 40.9 percent and pricing transparency at 38.3 percent. A review could receive more than one label, so these shares should not be summed.
In this study, what share of reviews was classified as mainly about the veterinary medicine itself?
In the main-reason classification, 28.1 percent of the 9,149 labelled reviews were assigned to the veterinary medicine itself and 70.6 percent to how the business is run. At least one core-service label was assigned to 39.6 percent of the 9,154.
Does the corpus include emergency or specialty vocabulary?
Yes. After 24 businesses were excluded, including 6 listed as emergency veterinarian services and 9 specialty or emergency referral hospitals, 14.3 percent of the 9,154 reviews still match a keyword pattern for emergency, specialist, referral, oncology, neurology or after-hours vocabulary. That is a keyword match, not a finding that a review involved emergency care.
Among the businesses sampled, what share of all reviews are one or two stars?
Among the 227,903 lifetime reviews held by the 296 businesses sampled, 6.7 percent are one or two stars and 86.7 percent are five stars. That is a share of reviews, not of clients.
How closely did the models agree?
On the binary main-reason question, how the business is run against every other main reason, a second model agreed on 92.5 percent of 200 reviews, kappa 0.83. For the theme labels, process against not process, a separate second model agreed on 97.0 percent of 200 reviews, kappa 0.796. No sample was checked against human coding.
Does the study verify whether what reviewers wrote is true?
No. It reports what reviewers wrote and how often the model assigned each theme. No record, estimate or treatment was checked against practice records.
Related studies
- The negative dental reviews study
- The negative plumbing reviews study
- The negative electrician reviews study
- The negative law firm reviews study
- The negative auto repair reviews study
- The HVAC profit benchmark review
- The negative pest control reviews study
- The negative moving company reviews study
- The negative towing company reviews study
- The negative property management reviews study
- The negative gym reviews study