Methodology
This page describes the road from reviews on Google Maps to a finished report: where the data comes from, how the numbers, problem cards and quotes are produced, what the language model does along the way and what plain code does - and where and why the method can be wrong. We write it so the result can be judged, not merely accepted.
Version of 24 September 2026
Where the data comes from
The source is public reviews of your business published on Google Maps - the same ones anyone visiting your listing can see. We do not buy data from anyone else and we do not run surveys.
Google does not confirm that a review's author was actually your guest or customer. Set the figures in this report against your own data - bookings, surveys, complaints - rather than basing decisions on a single source.
Collection is carried out by an external supplier named in our privacy policy. How far back we reach depends on the plan: a free account stores the most recent reviews, paid plans go deeper into the past.
Collection is a snapshot taken on a specific day. If an author later edits or deletes their review, the report still describes the state at the moment of collection. Dates and star ratings are taken exactly as Google publishes them - we neither correct nor fill them in.
The free report sample is a deliberate special case: it stands on the 50 most recent reviews of the business rather than on the whole body. The document says so explicitly on its first page ("Basis: 50 most recent of N reviews"), and sections that need a longer history - historical threads, for instance - simply do not appear instead of standing empty.
The QUALITATIVE findings - problem cards, strengths, quotes, narrative - are built on a SAMPLE of reviews rather than the whole body: up to 500 reviews go to the model, drawn proportionally to sentiment (in a network report: up to 100 per venue and up to 500 across the whole network, with a guaranteed minimum for every venue). For a large venue that is a fraction of the material, and the document says so explicitly in its method note, giving both numbers. The FIGURES - average rating, sentiment shares, star distribution, review counts - are computed from the WHOLE material covered by the report, not from the sample.
There are businesses for which we will not produce a report. If a venue has fewer than 3 Google reviews, we refuse the order before any payment - such a small body of material yields no findings about the business, and the resulting document would speak only about missing data. We check the same condition a second time after the reviews are downloaded, this time against their text: we need at least 3 reviews with text and 400 characters of that text in total, because a rating without a single word gives nothing to read. For an order covering several businesses the condition applies to their total, and a business with fewer reviews stays in the report and is described in it.
How the numbers are produced
Every review is assessed separately. A language model reads it and assigns a sentiment - praise, complaint or neutral - and the area it concerns, for example service, prices or cleanliness. One review may concern several areas at once, and then it counts towards each of them.
A separate step tidies up the area names: variants of the same topic - say, “service” and “customer service” - are merged, so one problem does not fall apart into two smaller counters.
Reviews are not translated before analysis. The model reads each one in the language it was written in, and the classification result goes into one shared vocabulary of areas - so a complaint in English and a complaint in Polish about the same thing land in the same counter. Translation enters only at the quote stage of the finished report - described in the section on quotes.
Reviews consisting of a star rating alone, with no text, count fully towards the number of reviews and the average rating, but they do not build problem cards or quotes - there is no content in them to read. That is why the report also shows the average separately for reviews with text: the two averages can differ, and that difference is information in itself.
Every number in the report - shares, complaint counts, charts - is a sum of those individual assessments. None of them is an estimate or a rounding: each can be broken back down into specific reviews.
Praise and complaint about one area are mutually exclusive within a single review: a generally positive review that complains about cleanliness counts for cleanliness as a complaint, not as praise. The strength card and the weakness card for the same area use the same rule, so they show the same number of praises.
The “stronger”, “weaker” and “similar” badges next to competitors are measured against the business’s Google rating over the whole history of its listing, not against the average of the analysed corpus - the two numbers can differ, especially when the report covers only the most recent reviews. The rating the badges were measured against is printed under the competitor table; the threshold is 0.1 of a star.
Owner replies are counted separately: the supplier returns the Google Maps reply in the same record as the review, and the report states how many reviews - and how many negative ones - have no reply. The figure covers only reviews carrying a measurement stamp; reviews fetched before 22 September 2026 have no such stamp, and the report then says “not measured” instead of giving a zero. The recommendation to reply to reviews is produced from this figure alone - no measurement, no recommendation. Our reply suggestions generated in the dashboard are not this figure: they say what could be written back, not whether the owner did. A recommendation appears only from 10 measured reviews - a prioritised recommendation based on one or two reviews would claim more than is known. The measured reviews are usually the newest, because measurement comes with each new fetch, so the figure describes current replying, not the profile's whole history; the report says so next to the figure.
Current problems, incidents and historical threads
The report does not throw all complaints into one bag. Complaints that recur and are recent become problem cards. Scarce signals, outweighed by praise of the same area, go to the “Isolated incidents” section. Complaints whose newest trace is old, while the area has since improved, land in “Historical threads - no recent signals”.
The same subject of complaint cannot stand in two states within one document. When a thread in the historical section describes the same matter as a card of a current problem - including when it is worded differently - the historical entry disappears, because a claim of silence is refuted by live evidence. The current card stays untouched. So does a single event in the incidents section - unless it is a second description of the same matter as the card (paragraph below): it then carries only what the card already carries.
Recency is measured against the end of the analysed date range, not the day the document was generated - a report covering last year judges recency as it looked at the end of that year. Complaint shares are calculated within the area they concern, not against the whole pool of reviews.
A card's weight - “Critical”, “Important”, “Watch” - does not follow from the raw number of complaints. It rises with the share of complaints among reviews about the area, the recency of the latest signals and the kind of area: topics touching safety or hygiene weigh more. That is why a card with fewer complaints can be heavier than one with more - the document points this out next to the cards as well.
One card is exempt from these thresholds: the card with the LIFE / HEALTH badge. It gathers reviews in which the language model - reading each candidate complaint separately - recognised an event threatening a guest's health or life: poisoning, injury, infection, insects, mould, fire, flooding, physical aggression, no help in an emergency or an unsafe structure. The code selects candidates (reviews with a complaint fragment the classifier assigned to health, safety, cleanliness, product quality and food, or - when a review has no fragments - a negative review tagged with such an area; a candidate's text is capped at 320 characters), the model judges, and the code enforces: the card is critical from the first report, stands first, and its figures count only confirmed reviews. Cards, incidents and historical threads about the same event are absorbed, so the same event does not appear twice under different weights. Why a verdict on the text rather than the classifier's categories alone: measured on 22 September 2026 on 45 random complaints from the areas “Health” and “Safety” - 28 concerned a guest's health or life, 17 did not (animal welfare, theft, a camera in the changing room, face masks); with that error rate a red badge would cost credibility instead of building it. The verdict can still err either way; every report on this card needs checking on your side. Only a record whose ALL quotes stand verbatim in a review recognised as a threat and describe the event itself, not another complaint from that review, counts as the same event; the category “Health” alone is not a reason. Such a record disappears when it has no more complaints than the LIFE / HEALTH card; a larger one stays with its own figures and its recommendations move under the LIFE / HEALTH card - the quote may then appear under both cards. A card with dozens of complaints about something else stays untouched. The first recommendation under this card is always the same: check every report in your own records; recommendations from the analysis follow it. The card becomes the biggest risk on the cover only when its latest report falls within the 12 months before the end of the report period - older reports stay on the card, but the cover speaks about a current problem. When the check did not cover all candidate reviews (time limit or model error), the PDF note says so plainly, and recommendations under health and safety cards stay in the document even with the status “Watch”.
The same problem does not appear in the document both as a current card and as an isolated incident or a historical thread - the document would then say two opposite things about one matter (“ongoing” and “faded”). The code picks the candidates (entries from the same category), and whether they concern the same subject of complaint is decided by the language model over the title and description of both entries: a shared category is not enough, because one area can hold several different problems. An entry that carries something the card lacks is not put to that verdict: an event marked as dangerous, an event from a location of the chain outside the card's evidence, or an entry newer than the date of the last report that the card gives - both entries then stay, because removing one would take the most recent signal about the matter out of the document. An entry judged to be the same problem is dropped, the current card stays with its numbers, and the recommendations that pointed to the entry move under the card and follow its rules - under a card with the status “Watch” they are not printed. When in doubt, both entries stay.
The section “What changed since the previous report?” shows new, persisting and closed problems and the average of new reviews only when at least 10 reviews have come in since the previous report. With fewer, differences between the problem lists would come from the analysis itself, not from customers - only the assessment of whether earlier recommendations worked remains. This is a different section from “What changed since the previous period?”, which compares two equal, consecutive periods of reviews; its dated items under “Sudden changes” are clusters of 1-2 star reviews within a short window, clearly above the usual pace of that location - clusters of 5-star ratings are not listed there, and the same cluster of one-star and low ratings counts once. When the period comparison has no items, the section is not printed.
With a LIFE / HEALTH card, the row “Top Risk” in the summary on the first page points to this card - the number of reports and a reference to it - and “Act Now” is the recommendation to check every report under this card (from the immediate and short-term recommendations). The row “Estimated impact” from the analysis is then omitted, because it described a different action. The card is critical and stands first, so the summary cannot name something else as the biggest risk. Its date of the latest report comes from the most recent confirmed event, even when that event is not quoted.
The same quote - in the same wording - stands in the document in one place. A shorter excerpt of a review quoted at greater length elsewhere is not treated as a repetition, because it usually carries a different thread of the same review; it can, however, be the same complaint under two entries. When the same customer sentence would go under two entries, it stays with the earlier one in this order: problem cards, incidents, historical threads, strengths - a mixed sentence (“tasty, but cold”) is evidence of a complaint. When a quote stands both under a problem card and under an incident or a historical thread, it stays with the incident or thread, because these usually have one piece of evidence and the card has a whole pool. Exceptions: an entry whose quotes are all repetitions keeps one - the last - because an entry without evidence would be worse than a repetition; a quote that is the only evidence of a chain location listed next to the entry also stays; and finally the LIFE / HEALTH card, under which a quote about a threat may stand next to the card it was taken from.
The formulas the report stands on
The complaint share in an area is calculated from reviews that speak about that area: complaint share = complaints about the area ÷ reviews about the area. The denominator is not the whole pool of reviews - a complaint about parking does not dissolve in a hundred compliments about the kitchen.
Sentiment shares on the ring chart are calculated with the largest-remainder method so they add up to exactly 100%, and the net sentiment is derived from the same shares: net sentiment = % praise - % complaints. One definition for the chart and for the text means two numbers about the same thing cannot drift apart even by a point.
Recency of a signal is a difference of dates: signal age = end of the analysed range - date of the newest complaint in a thread. Card weights and the historical-threads section use the same difference.
A card's weight is a threshold rule, not an average. Simplified: a card receives the “Critical” badge when complaints in the area are both numerous and recent, or when a high-risk area - safety, hygiene - carries a recent, material signal. The count thresholds scale with the size of the material: with a few thousand reviews the bar hangs higher than with a hundred, so a handful of complaints in a large venue does not get a red label. The exact threshold values belong to the code and change with it, which is why we do not transcribe them here.
Where the quotes come from
Quotes are reproduced verbatim, exactly as customers wrote them, typos included. We do not shorten them in a way that changes their meaning and we do not join sentences from different reviews.
Each quote carries its publication date, and in a report for a chain, the location it concerns. That way every quote can be found on Google Maps and checked.
A quote written in a language other than the report's is shown in translation, with a small “(translated)” note next to it. The quote is selected and its date matched always on the original, word for word - only the finally chosen version is translated, and the translation is done by the language model.
Choosing a quote for a card is the model's decision, so a separate check verifies after the fact that the quote is on the topic of the card it stands under. Matching a quote to a location and date is done by code - a simple lookup against the source review, with no model involved.
What the model does and what the code does
Two models of different strength work on a report: a faster one reads every review separately and classifies it, a stronger one writes the narrative content - the prose, the problem descriptions and the proposed actions. The model handles everything that requires understanding a sentence.
The code handles everything that can be checked without understanding: numbers, dates, summing the classifications, matching a quote to a location, keeping sections consistent with one another. To the code, the model's output is material to verify, not an oracle.
A finished report passes a set of checks before we hand it over. Among other things we check: whether a card's title describes one subject of complaint rather than several at once; whether the evidence under a card is on that card's topic; whether two cards describe the same problem; whether the text contains non-existent words, that is, the model's typos; and whether every area with a visible problem has its card. Some of these checks are plain code, some are the model's second judgement over the finished text. A check that is not certain leaves the text alone rather than forcing a correction; a report that fails the checks is generated anew.
A chain report versus single-location reports
A chain report is not a paste-up of location reports. All of the chain's reviews are analysed as a single body: a problem visible in several locations becomes one card marked with how many locations it concerns, and a problem of a single location goes to the part about local matters. Every quote carries the location it comes from.
That is why numbers in a chain report can differ from numbers in single-location reports - and it is not a contradiction. Topics in a chain report are broader: one chain card can span several narrower local cards, so its counter can be larger than their sum while there are fewer cards. Both perspectives describe the same material at a different grain.
A sentence in the description of a card, thread, incident or strength that names a location of the chain not covered by that entry's evidence - neither the list of locations at the entry nor the quote labels - is removed in full; an entry whose whole description concerns such a location stays unchanged. A recommendation is judged differently, because it does not stand next to quotes but directs action: a location named in a recommendation must have a complaint in the document from the same area as the entry the recommendation points to - in any section, including local problems - or praise from that area among the strengths - which counts only when the recommendation also names a location with a complaint in that area; the praised location is then treated as a model - the code does not recognise a location's role in the sentence, it only checks this proximity. A location without such support that stands in a list of locations next to supported ones is removed from the list together with its conjunction, and the rest of the recommendation stays - provided it really is a list shaped “A, B and C” (commas first, then “and” - no comma after the conjunction): at least three locations, or two joined by a conjunction at the end of the sentence with no plural noun right before them - and provided the word right before the list is a preposition of place (“in”, “at”, “for”) or a location noun (“the location”, “the property”), the list ends the sentence, and the sentence contains no numeral (other than one of time, money, percentage or number of times, such as “every two weeks”) and no “each” or “all”. Otherwise the rest of the sentence, a word before the list (“except”, “between”) or a number (“both”, “in 3 locations”) could refer to the very item being removed. In every other case the whole recommendation is dropped, because a comma between two names can be the boundary of two clauses, and removing a name would change the meaning. A recommendation without such footing is not printed at all, rather than in part, because removing one sentence could take away the instruction itself; only a closing “… at location X” is removed from the title. Recommendations that point to no entry are judged by the recommendation footing check. In the table of actions for each location, the sentence about maintaining the standard does not stand next to a location that has a critical or important card in the document - it is replaced by a reference to that card with this location's number of complaints, when the card gives it, and the recommendation under it. A location with reports on the LIFE / HEALTH card gets a reference to it also when the action in the table was written by the analysis - that action follows the reference. A recommendation goes into a location's row only when it does not name another location of the chain. For a location whose most recent review is older than 12 months before the end of the report period, the document gives that review's date in the table of locations, in the best and worst location box and in the table of actions - the assessment of such a location describes the state of more than a year ago.
How this differs from pasting reviews into an AI chat
The most common question about this product is: why pay, when reviews can be pasted into a free AI chat. The first difference is mundane: the chat does not have the data. Google Maps does not hand reviews over conveniently in bulk - collecting them across the whole analysed range, with dates and star ratings, is work in itself, and everything further is computed on that set, not on the fragment that fits into a conversation.
The second difference is how counting works. A model reading thousands of reviews at once does not count - it estimates. Asked how many complaints concern service, it will give a number nobody can check. Here every review is classified separately, and every number in the report is a code-computed sum of those individual verdicts - and each can be broken back down into specific reviews.
The third difference is control. A chat's answer is a first draft that nobody checks. This report passes the set of checks described above - evidence under cards, duplicates, area coverage, non-existent words - and the consistency of numbers between sections is guarded by code, not by the model's memory. On top of that come rules applied the same way to every business: significance classes, thresholds scaled to the size of the material, recency measured against the date range.
Honestly: with a handful of reviews the difference is small - a good chat will summarise a few dozen pasted reviews sensibly. The difference grows with the material: with hundreds or thousands of reviews, several locations and a long date range, a single-pass summary stops being countable and checkable - and countability and checkability are exactly what this document answers for.
Where the data can mislead
Reviews describe what customers chose to write. Satisfied customers write less often than disappointed ones, so a picture drawn from reviews is not the same as a picture drawn from a survey of all customers. The report measures the voice of those who write - and only theirs.
When there are few reviews, a single statement weighs a great deal. The report marks such places with a note about sample size, but the note does not change the nature of the thing: it is still a conclusion drawn from thin material, and one new review can overturn it.
The historical-threads section reasons from silence: since there are no recent complaints, the problem has probably ceased. Silence, however, is not proof - customers may have stopped writing for other reasons, and the problem may persist. The word “probably” next to such threads should be read literally.
Where the automation can err
For ambiguous reviews - praising one thing and criticising another - the classification can be wrong, and the review lands in the wrong area. Every number in the report inherits the quality of that classification: in a counter built from hundreds of reviews single mistakes drown, but in an area with a handful of reviews one mistake visibly shifts the picture.
The narrative is written by a model, so writing faults happen: a card's title can stress a minority of the issues it covers, a description can go a step further than the evidence reaches, and the text can contain a typo forming a word that does not exist. The checks described above catch a part of such faults - not all of them.
The checks themselves are fallible too, in both directions. We tuned them cautiously: a check without certainty leaves the text alone, because a forced correction would also remove true content. The price of that caution is that an occasional fault can slip through the sieve - for example, two cards about very similar problems can occasionally stand side by side.
The split into systemic problems, incidents and historical threads rests on thresholds of signal count and recency, and every threshold divides something by a hair. A complaint that is rare but serious can end up among the incidents even though you would judge it heavier. That is why the incidents section shows such signals openly instead of hiding them.
The complaint count on a card comes from the scope check: a language model marks which complaints speak about the topic in the card's title. It does this three times, independently, and a complaint counts towards the card when at least two of the three readings mark it - a single reading could give the same card a dozen complaints one time and over a hundred another, on the same reviews. Sometimes a problem card stands in the report without its own complaint count and without a weight badge - it is then marked “Weight: not measured”. This happens when the scope check found no reviews that unambiguously speak about the card's narrow topic, or when it could not check a card added at the end of the analysis. Rather than substituting the numbers of the whole, broader category - which would describe something other than what the title claims - the report shows the category numbers with their population named and openly admits it did not measure the topic's coverage. Read such a card as a signal for your own verification, not as a measured conclusion. An action based on such a card is printed without a cost estimate and without high priority - spending on a problem whose scale the report did not measure gets no cost bracket in the report.
Very rarely the graphic layer of the document cannot be composed - with an unusually long name or an unusual set of data. The report then comes out in plain form: full content, all figures and quotes, without charts or coloured cards. The document says so in its opening paragraph and gives an address for requesting the full version. We would rather deliver a report without its graphic layer than deliver none.
In a network report, a problem card usually covers several locations at once. Next to each name we then give that location’s own complaint count and its number of reviews with complaints: “Willa Hyrny 13 of 84” means thirteen of that location’s eighty-four reviews with complaints concern the card’s topic. The order in that line follows the share, not the raw count - a smaller location may have fewer complaints yet the worst proportion, and it then comes first. The breakdown appears only when the parts add up to the figure printed under the card and when it can be computed for every location listed; otherwise the card shows a single combined figure, which must not be attributed to each location separately.
The list of locations on a card - the “Affects” line - follows the evidence, not the model’s declaration alone. When the scope check finds not a single review about the card’s topic at a listed location, that name is removed from the list; the card’s figures do not change, because that location contributed nothing to them. One exception: a location that the card’s description mentions by name stays on the list, and the card then shows a single combined figure without a breakdown. A location missing from the list does not mean the problem is absent there - only that the reviews covered by the report gave us no evidence of it. A card added by the area-coverage check, that is when an area with a visible problem received no card of its own, counts complaints only in the locations it lists, not across the whole network.
The recommendations in the report are drawn from reviews alone and know nothing of your costs, rosters, contracts or booking data. They are hypotheses to check, not the conclusion of an operational audit: the report can show how many customers write about a problem and since when, but it does not know whether the cause lies where it suggests, nor whether the fix fits your budget. Verify them on your side before any staffing, pricing, operational or investment decision.
A recommendation must stand under a card that justifies action. A card with the status “Watch” - a signal below the volume, share or recency threshold - is shown with its figures, but no recommendation is printed under it: the document cannot tell you to watch and to spend in the same place. The same applies to a recommendation that points to no card, incident or historical thread in this document - it is dropped from the report instead of receiving a caveat. The exception is an explicitly maintenance-type recommendation in a report with no current weaknesses. Measured on 586 saved reports (22 September 2026): the rule removes 1328 of 2785 recommendations, 889 of which stood under “Watch” cards; 126 reports are left with no recommendation at all, and the document says so in its note on method. Most of them are reports from before 4 August 2026, when recommendations did not yet have to point at a card: among reports since that day, 14 of 184 are left without recommendations. The note states how many recommendations were dropped, and when all of them were, the recommendations chapter stays in the document with one sentence saying so instead of disappearing. A “maintenance” mark does not save a recommendation that points at a card absent from the document.
Reviews flagged for verification
Beyond the content, we check one thing against the calendar: whether reviews appeared in an unnatural cluster. We divide the company's history into non-overlapping two-week windows and count how many reviews fall into each. The threshold is not one number for everyone - we set it separately for each company from a distribution fitted to its own history, so that ordinary fluctuations in traffic rarely cross it. We count clusters separately in two series of low ratings - in one-star reviews alone and in one- and two-star reviews together - because a cluster of the lowest ratings means something different from an ordinary bad week. For half the companies in our set the threshold falls at three such reviews in two weeks, and for a company with denser traffic it can be almost four times higher. Three- and four-star ratings do not enter this calculation: we measured this across our whole set, and their presence nearly doubled the number of false indications without improving detection of real clusters. For most of the companies we examined the arithmetic alone would put the threshold lower still - but we never go below three reviews, because two reviews in two weeks is an ordinary week, not a pattern. The check does not run when the history is shorter than roughly three months, because a shorter one cannot establish what is usual for a company. Since 6 September 2026 the check runs with every review synchronisation, and its result is shown in the company panel on a separate tab - together with which reviews it concerns. No report needs to be generated for this.
A cluster on its own settles nothing and is not enough to flag a review. It is passed to the model as circumstantial evidence, and the model may flag a review only when it finds at least two independent grounds; in case of doubt it is instructed not to flag. Harsh but specific and credible criticism of a real visit is not grounds for flagging.
A flag is information, not a verdict. We do not rule that a review is untrue, and we do not assess the person who wrote it. Flagged reviews REMAIN in every figure in the report - the average rating, the star distribution and the sentiment shares - precisely because the flag is a prompt to check, not a decision. Whether a review breaches Google Maps policies is decided by Google alone.
We do not detect every coordinated action, and we say so plainly. The check covers BOTH families of reviews - low ones and five-star ones - because what marks a burst as unnatural is the rhythm in which reviews arrive, not whether they are kind. The two lists in this service have different scopes, and the difference is worth naming: the list of clusters in the panel covers only 1- and 2-star ratings, because clusters are counted on those alone, while the list of flagged reviews inside the report itself covers ratings from 1 to 4 stars. In the study itself, only a period of 1- and 2-star ratings can be a marked period: an unusual influx of praise is not marked there as a period, because nothing follows from it for you - we do not report five-star reviews. We do keep counting them, and that works the other way round: if more of the remaining ratings arrived at the same time, the period of complaints is explained by ordinary traffic at the venue and we point to no candidates from it. We do not challenge praise on your own profile: we tell you about it separately, without pointing to individual reviews, because Google penalises businesses for bought reviews even when someone else ordered them. We also see action spread over time less well: a cluster stretched across several windows raises what counts as usual for the company and so partly hides within it. The absence of a flag therefore does not mean that nothing happened. Since 8 September 2026 the timing check itself is a separate, paid Review cluster study - it can be ordered without an account or added to a one-time report. In the Service this service is presented under the name “Review verification against Google's rules”. The result of the study is a page listing the marked periods and the reviews that fall into them; the report gives only the number of periods. In the study, a language model additionally reads the text of every review of the business - regardless of the star rating and with no cap on their number - and rules on one thing only: whether the statement itself breaks the content policies that apply on Google Maps - for example contains obscenity, advertising or an attack on a specific person. The star rating alone never qualifies a review for the list - only its wording does. One limit works the other way: we do NOT put five-star reviews on the list to report, even when they break the content policies. We advise against reporting them: removing praise would lower your rating, working against the very purpose you ordered this service for. We still read them, so that the selection is not driven by negativity alone. The model reads every review twice and independently, and a breach flagged in either reading goes on the list - in our measurement a single reading missed between a few and thirty per cent of the breaches that the second reading found (on average around one in six). Obscene words are additionally guarded by a dictionary in our code: when neither reading of the model flags a vulgarity that is in the text, the review still goes on the list with that word as the quote - in our measurement the model missed nearly two such words in five. A review from a marked period without a breach of its own is shown as a candidate for the client's decision only when the period is not explained by the venue's general traffic (no rise in the other ratings at the same time) and the text carries no specifics of a visit; the decision rests with the client. The report justification is produced in two stages. The facts - the breach category, the verbatim fragment of the review, the dates and numbers of the period and the quoted Google Maps rule - are set by our code. A language model arranges them into sentences. The result is then checked automatically: we reject any text containing a number, a quote or a claim from outside the set of facts, and substitute a ready-made template instead. Neither the model nor the code writes a sentence judging the truth of a review or its author - Google rules on a breach of its policy, not on the truth of a statement. In our measurement on the whole corpus (8 September 2026, over twelve thousand reviews with text) fewer than one review in a hundred had such a breach, so the list is often short or empty.
What this report does not do
It does not cover reviews from services other than Google Maps, nor events customers did not write about.
It is not legal, tax or investment advice. Cost ranges are indicative and serve to set the order of work, not to order services without your own judgement.
The role of artificial intelligence
The narrative content of the report - the prose, the problem descriptions and the proposed actions - is produced with the assistance of a language model. We disclose this in every report.
PDF files additionally carry a machine-readable marking indicating that the content was generated by an artificial intelligence system. This corresponds to Article 50(2) of Regulation (EU) 2024/1689 - the EU Artificial Intelligence Act.
Purchase, payment and complaint rules are set out in the terms of service. This page neither replaces nor changes them.