AI Search Responses Meaningfully Shift Brand Position
Kotler and Keller (2003) define brand positioning as ‘the act of designing the companys offering and image to occupy a distinctive place in the mind of the target market.’
The Branding Journal (Marion, n.d.), explains that ‘brand positioning describes how a brand is different from its competitors and where, or how, it sits in customers’ minds. A brand positioning strategy therefore involves creating brand associations in customers’ minds to make them perceive the brand in a specific way.’
Brands have traditionally controlled the format, structure and messaging, customers read, see or hear about their market position. Advertisements across radio, television and social media typically undergo formal review and approval to ensure they communicate the intended brand positioning.
In organic search, brands maintain this control through elements such as the URL slug, page title and meta description displayed in search results. When customers visit the brand’s website, the brand also controls the page content and how its positioning is presented.
AI Search is the first medium where brands do not have full control over their brand positioning. The AI Search response may mention a brand, but will also independently explain,
- how a brand is different from its competitors,
- the benefits of the brands product and services to the customer,
- the reasons why the brands product and services may not align with the customers buying intent
This effectively interjects the brands message and could intervene to shift the brands position in the mind of the target market.
Somantra studies the impact of AI Search intervention to brand positioning in AI responses. We looked closely at 135 different search queries across multiple insurance product categories. Our main finding is that your brand may still appear if a AI search query is re-run by adding/removing/changing a word, but in the new AI response it could lose its reason to be chosen by the customer.
We added decision words like cheapest, safest, or most trusted to the ChatGPT search prompts, selected the responses which had a table with brand comparisons for our study..
The perturbation in the search query did not just change the order in which the brands were mentioned by ChatGPT, but it also shifted its narrative about each individual brand.
In one Home Insurance search, we added the word safest. The impact of adding this high intent word has different impact on AAMI Vs QBE.
| Query | AAM | QBE |
|---|---|---|
| home insurance providers in Australia | Large mainstream insurer; home & contents | Major Australian insurer with home & contents products |
| safest home insurance providers in Australia | Consistently strong value/coverage results and particularly strong Queensland performance. Canstar named it a 2025 Outstanding Value winner nationally and in QLD. | Good option if comprehensive/flood cover is particularly important; Mozo’s 2026 guide selected QBE for flood cover. |
Both brands remained on the list of results but for AAMI its positioning shifted meaningfully to raise brand consideration with the customer, while for QBE it was presented simply as an option.
There can be a big gap between just showing up in an AI search response vs. showing up with the right message to win against competition. Most brands are measuring their mentions in AI Search but not whether this mention is supported by messaging which improves brand consideration by the customer.
The Executive Readout: 4 Big Shifts for Your Team
Our research highlights three main decisions for marketing, SEO, and brand teams.
- Measure the message, not just the mention. Showing up is simply not enough anymore. A simple mention count does not tell the whole story. It cannot show if an AI highlights your brand for price, trust, safety, or even a bad warning.
- Watch how customers make decisions, not just single prompts. Customers express their search intent using many different words. Each new word creates a totally different competitive space in AI Search for brands.
- AEO/GEO strategy should focus on the brand getting chosen, not just mentioned Customer will see multiple brands in the AI Search response with AI shaping the message for each brand. AEO/GEO strategy should focus on influencing AI to shape this to favour the brand.
Perturbation testing is a simple but powerful idea. It means you change one specific part of a search prompt. You keep everything else exactly the same. Then, you measure how the AI answer changes.
In the current study, ChatGPT responses were collected from a single run on the queries and a limited set of word perturbations. The next step is a comprehensive test with expanded set of word perturbations and ChatGPT responses collected across multiple runs. You can find more details about our research boundaries in the appendix at the end of this post.
1. How Changing One Word Rewrites the Narrative
Open Data, Code, and How to Read the Metrics
The full code and dataset for this study are public. You can inspect the queries, reproduce every figure and table, and audit the method in our GitHub repository: Somantra/AEO-research (perturbation_chatgpt_tables).
Three metrics appear throughout this section, each reported on a 0 to 1 scale:
- Word shift measures how much of the visible wording changed, using Jaccard distance. A score of 0 means the answer reused the same words, and a score of 1 means it used a completely different vocabulary to describe the brand.
- Semantic distance (meaning shift) measures how far the underlying meaning moved, using embedding similarity rather than exact words. A score near 0 means the framing stayed similar, and a score near 1 means the brand’s rationale was reframed entirely.
- Net sentiment shift (tone shift) measures movement toward positive wording minus movement toward negative wording. A positive value means the answer became more favourable to the brand, and a negative value means it became more cautionary.
Figure: The single modifier word that caused the largest shift in meaning for the top volatile brands.
Price as an Intent Changes Everything
Adding words like cheapest or most affordable caused the biggest shifts. These words changed how AI described brands the most.
We saw an average 0.97 word shift on a 0 to 1 scale. A score of 1 means the AI used completely different words compared to the original search associated with the brand..
Figure 2. Average vocabulary shift by modifier. The worst category is not plotted because it only has one observation. It is still in the table below and is a focus for deliberate future tests.
When a customer asks about price, AI Search is looking for brand facts along different words. It might look for low premiums, local quotes and special discounts. It might overlay this with rules about excess fees; in some cases it might even show a clear warning that a provider is not a cheap option.
Because of this deep search for new facts, your brand might stay visible. However the explanation provided by AI for your brand can shift towards your aspired position or significantly away from it.
The Shifts Are Big, but They Vary
Across all 135 observations, the median word shift was 0.94. The mean shift was 0.90.
The median is a bit higher because a few smaller changes pull the average down. Simply put, a typical changed prompt rewrote almost all of the visible text. But the exact impact was not exactly the same every single time.
We also measured semantic distance. This checks how much the actual meaning changed, not just the exact words. A value of 0 means the meaning stayed very similar. A value of 1 means the framing became completely different. The median semantic distance was 0.53.
Three patterns stood out very clearly in the data:
- Price drove the biggest text rewrites. Words like cheapest and most affordable changed the vocabulary the most.
- Best changed the reasoning. When we used the word best, the median semantic distance was high at 0.60.
- Trust changed the emotion. The phrase most trusted created the biggest shift toward positive feeling. It moved sentiment up by positive 0.50 across 16 observations.
Figure: Average shift towards positive vs. negative sentiment language by modifier.
| Modifier | Observations | Mean Word Shift | Observed Word Range | Median Semantic Distance | Middle 50 percent of Semantic Distance | Mean Net Sentiment Shift |
|---|---|---|---|---|---|---|
| Most popular | 20 | 0.80 | 0.38 to 1.00 | 0.48 | 0.31 to 0.51 | positive 0.09 |
| Easiest | 19 | 0.92 | 0.71 to 1.00 | 0.54 | 0.49 to 0.61 | positive 0.15 |
| Best | 17 | 0.94 | 0.82 to 1.00 | 0.60 | 0.53 to 0.67 | positive 0.35 |
| Cheapest | 17 | 0.97 | 0.83 to 1.00 | 0.63 | 0.54 to 0.72 | positive 0.19 |
| Most trusted | 16 | 0.83 | 0.60 to 1.00 | 0.47 | 0.43 to 0.52 | positive 0.50 |
| Most affordable | 13 | 0.97 | 0.88 to 1.00 | 0.60 | 0.52 to 0.65 | positive 0.10 |
| Most reliable | 13 | 0.90 | 0.60 to 1.00 | 0.51 | 0.41 to 0.66 | positive 0.27 |
| Safest | 13 | 0.89 | 0.71 to 1.00 | 0.52 | 0.41 to 0.65 | positive 0.37 |
| Premium | 6 | 0.90 | 0.83 to 1.00 | 0.40 | 0.32 to 0.49 | positive 0.08 |
| Worst | 1 | 0.94 | 0.94 | 0.76 | Not Applicable | negative 0.21 |
Table 1. Two decimal results with sample sizes and spread. Net sentiment shift means movement towards positive language minus movement towards negative language. The single worst observation is descriptive only.
Category Matters
Different types of insurance products require different types of proof.
Home Insurance had a very high median word shift of 0.98. It also had a mean semantic distance of 0.59 across 46 observations. Motorcycle Insurance gave us 31 observations. Roadside Assistance gave us 34. Travel Insurance gave us 24.
Figure 3. Home Insurance recorded the largest observed movement in this sample. Word and semantic measures use different scales and should be read separately.
The practical lessons are
- Different products, services require different evidence for AI Search to choose and recommend a brand.
- A single AI visibility score(e.g. mentions/citations) across all categories fails to explain the decision lens that AI Search applies to generate its answer.
2. What Perturbation Means in AI Search
A perturbation is a controlled and minimal change to a query. For example adding a single modifier while holding the rest of the query constant. It estimates answer engine sensitivity to wording.
The method is a practical version of a minimal pair experiment. Instead of asking whether a brand appears in a broad set of unrelated questions. It asks exactly what happens when one single decision driving element changes.
The corpus reconstructs single modifier pairs from historical query logs. So its findings show observed sensitivity and association. Section 6 sets out the repeated query controls needed for a fully controlled experiment.
| Term | Meaning | Why it matters |
|---|---|---|
| Base query | The original prompt, such as “home insurance providers in Australia.” | Establishes the comparison point. |
| Perturbation modifier | The one deliberate change, such as adding “safest” or “cheapest.” | Tests a specific decision lens. |
| Minimal pair | The base and perturbed query considered together. | Keeps the comparison interpretable. |
| Response shift | The observed change in a brand wording role sources or inclusion. | Captures change that a simple mention count misses. |
| Word shift | Change in visible wording measured here with Jaccard distance. | Shows how much the positioning language was rewritten. |
| Meaning shift | Change in semantic framing measured with embedding similarity. | Distinguishes a real rationale change from simple adjective swapping. |
| Tone shift | Change in the positive or negative valence of the positioning phrase. | Shows whether the answer becomes more favourable or more cautionary. |
| Query neighbourhood | The cluster of nearby variants around the same customer need. Like price trust safety convenience and audience fit. | Defines the space a brand must be ready to answer within. |
| Noise floor | The movement seen when the exact same prompt is repeated multiple times. | Separates modifier effects from normal answer engine variability. |
Perturbation is not simply keyword research with a new name. In a standard conventional search results page an added word may change which URLs actually rank. In an answer engine it can also completely change the argument the model constructs. It changes which specific attributes it foregrounds which sources it cites and what exact role it gives a brand.
3. What the Study Observed
Price Modifiers Triggered the Largest Rewrites
The study started with 4,445 ChatGPT responses to Australian insurance shopping queries. These all contained brand comparison tables. We kept only single modifier pairs where the same brand appeared in both answers. We excluded multiple term rewrites completely. This left 135 clean observations across Home Insurance (46), Roadside Assistance (34), Motorcycle Insurance (31), and Travel Insurance (24).
Adding the word cheapest or most affordable produced the largest average(absolute) word shift movement. That score was 0.974. Best followed closely at 0.942 and easiest at 0.921. The practical meaning is not that price is the only relevant criterion but that a price request deeply changes the proof an answer engine must assemble. It looks for a low premium discount caveat regional offer or a reason a brand is not price led.
Meaning Changed Not Just the Adjectives
The sample had a word shift of 0.944. At the same time the median meaning shift was 0.534. Those are different measures for a good reason. An answer can definitely use new words while making the same basic case, or it can keep familiar words while assigning a brand a totally different comparative role.
Figure: Emotion vs. Meaning Shift mapping modifiers by semantic and emotional volatility.
Figure 2. Home Insurance showed the largest observed movement in this sample. Wording and semantic measures are shown separately because they use completely different scales.
Home Insurance had a median word shift of 0.983. It also had a mean meaning shift of 0.590 across 46 clean observations. This does not automatically establish that every single home insurance answer is more volatile than every other category. It simply shows why category specific perturbation testing is much more useful than a single undifferentiated AI visibility score for a brand.
A Modifier Can Change the Job a Brand Does in the Answer
In one specific matched Home Insurance pair ChatGPT described AAMI without a modifier as a Large mainstream insurer home and contents. With the word safest added the description became a highly detailed value and coverage rationale completely tied to Queensland performance and a major award.
Figure 3. One matched observation. It records a direct difference in the observed response. It is not an endorsement or a stable permanent claim about the insurer.
This is the central reason to care very deeply about perturbation. The brand can definitely remain present while its positioning changes completely from broad availability to value trust safety affordability or a simple caveat. Simple presence alone would report no real change. But the actual customer facing answer has changed substantially.
4. How Perturbation Fits Into AI Search Measurement
Somantra uses two very complementary lenses for this. They answer totally different questions and must not be merged into one single metric.
| Lens | Temporal measurement | Perturbation measurement |
|---|---|---|
| Question | What changed across repeated market snapshots? | How sensitive is an answer to one small wording change? |
| Design | Same normalised queries observed in different months. | Minimal pair base and modified queries. |
| Best for | Brand inclusion source reweighting concentration and persistence. | Positioning resilience decision criteria and fragile query variants. |
| Control | Matched query cohort. | Minimal change plus repeated identical query controls. |
| Scope | Market level observation. | Local per query sensitivity. |
The companion temporal study(AI Search Visibility: Australian General Insurance Brands - May 2026 Report) found that ChatGPT brand inclusion and citation mix changed significantly even on fixed query cohorts run across months. That clearly establishes exactly why monitoring must go far beyond simple raw volume. Perturbation testing adds the crucial next layer. It helps identify the specific decision words under which a brand description evidence or comparative role is most likely to change.
Neither lens specifically identifies an engine internal algorithm or directly proves a commercial outcome on its own. Temporal analysis simply reports observed change over time. Perturbation analysis directly estimates local sensitivity to a highly controlled wording difference. Together they make all AI search measurement much more diagnostic and much more actionable for marketing teams.
5. Action Plan for Marketing and SEO Teams
1. Build Query Neighborhoods Around Customer Decisions
For every single high value base question test one carefully controlled variant at a time. A multiple word rewrite is a completely new prompt. A single perturbation changes one interpretable variable at a time.
- Price: cheapest affordable value
- Trust: trusted reliable reputable
- Protection: safest comprehensive best cover
- Ease: easiest simple claims easy to manage
- Fit: families young drivers regional customers trip type
2. Give Every Brand a Resilience Profile
Track everything from mentions,citations, observed positioning, word shift, semantic distance and sentiment movement by modifier. Always pair every result directly with its exact observation count. Avoid turning sparse data into arbitrary and misleading leaderboards.
3. Build Evidence for the Decision Lens
The main goal is absolutely not to publish a new page for every single adjective. It is simply to make price conditions coverage details awards claims evidence and customer fit information perfectly accurate. Make them specific and highly retrievable by AI systems.
4. Measure the First Screen
Always store both the generated text response and the actual rendered visual result in AI Search response. A simple brand mention that survives in text but surfaces below the first mobile viewport drives lower brand recall with the customer.
5. Establish the Noise Floor
Repeat exactly identical base queries. Completely randomize the execution order. Compare modifier driven movement directly against normal repeat variation. This turns simple directional evidence into a robust causal sensitivity estimate.
Also please note that ChatGPT advertising should definitely be added as a separate commercial exposure layer later. It should absolutely not be merged silently with any organic perturbation metrics.
Research Appendix
Evidence boundary: The 135 published observations were reconstructed from historical query logs. They were not assigned in a randomized experiment. They demonstrate observed sensitivity, not a guaranteed causal modifier effect. The corpus has no repeated identical query controls yet. So it cannot estimate the model natural response variability. The recent audits provide temporal context but contain no systematic perturbation grid. Mobile exposure and advertising are proposed measurement layers. They are not findings in the current dataset.
Full Observed Brand Profile
| Brand Label | Pairs | Modifiers | Median Framing Change | Largest Observed Change | Coverage |
|---|---|---|---|---|---|
| 1Cover | 5 | 5 | 0.54 | Most trusted (0.88) | Broader |
| AAMI | 13 | 9 | 0.58 | Most reliable (0.82) | Broader |
| AANT | 1 | 1 | 0.48 | Most popular (0.48) | Limited |
| AIG | 1 | 1 | 0.76 | Worst (0.76) | Limited |
| Allianz | 9 | 8 | 0.61 | Cheapest (0.72) | Broader |
| Apia | 2 | 2 | 0.55 | Cheapest (0.72) | Limited |
| Budget Direct | 15 | 6 | 0.53 | Easiest (0.67) | Broader |
| Cover More | 5 | 5 | 0.41 | Easiest (0.54) | Broader |
| GIO | 2 | 2 | 0.78 | Most affordable (0.79) | Limited |
| InsureandGo | 3 | 3 | 0.62 | Easiest (0.85) | Limited |
| National Motorcycle Insurance | 4 | 4 | 0.51 | Cheapest (0.61) | Limited |
| NRMA | 12 | 8 | 0.63 | Best (0.83) | Broader |
| NRMA Insurance | 3 | 3 | 0.65 | Most reliable (0.71) | Limited |
| QBE | 9 | 8 | 0.57 | Easiest (0.76) | Broader |
| RAA | 5 | 4 | 0.51 | Best (0.67) | Broader |
| RAC | 2 | 2 | 0.48 | Most popular (0.52) | Limited |
| RACQ | 9 | 7 | 0.50 | Most trusted (0.58) | Broader |
| RACT | 2 | 2 | 0.38 | Most popular (0.45) | Limited |
| RACV | 9 | 6 | 0.51 | Cheapest (0.75) | Broader |
| Shannons | 5 | 5 | 0.39 | Most reliable (0.45) | Broader |
| Suncorp | 3 | 3 | 0.57 | Safest (0.65) | Limited |
| Tick Travel Insurance | 1 | 1 | 0.45 | Most affordable (0.45) | Limited |
| Travel Insurance Direct | 1 | 1 | 0.38 | Safest (0.38) | Limited |
| World Nomads | 4 | 4 | 0.39 | Most trusted (0.41) | Limited |
| Youi | 10 | 9 | 0.65 | Best (0.75) | Broader |
Table 4. All observed brand labels. Broader means at least five pairs. It is not a quality rating. NRMA and NRMA Insurance remain separate source labels. They should be reconciled before external entity level benchmarking.
Method in Plain English
The source corpus contains 4,445 ChatGPT response files from Australian insurance shopping queries with comparison tables. We matched brand rows across related queries. Then we retained the 135 pairs where one modifier was added and the same brand appeared in both tables.
- Word shift: how little vocabulary the two positioning phrases shared.
- Semantic distance: how far their meanings moved. This is based on numerical language representations called embeddings.
- Net sentiment shift: movement towards positive wording minus movement towards negative wording.
- Matched temporal cohort: the same normalized questions observed in February, May, and July.
- Domain concentration score: whether citations were spread across many domains or clustered among fewer sources.
Values in the narrative are rounded to two decimals. Aggregate modifier and brand tables are retained with the Somantra research record.
Research source: Somantra Query Perturbation Study: ChatGPT Brand Positioning and Somantra AI Search Temporal Research. Australian insurance corpus. Perturbation analysis run 8 August 2026.
Search Query Perturbation Impact Analysis Tool
Frequently asked questions
What is query perturbation in AI search? +
A perturbation is a controlled, minimal change to a search prompt, such as adding a single modifier like 'safest' or 'cheapest' while keeping the rest of the query the same. It measures how sensitive an AI answer engine is to that one wording change, isolating the effect of a single decision word.
Can changing one word in an AI search prompt really reposition a brand? +
Yes. Across 135 single-modifier observations spanning 25 Australian insurance brands, the median change in visible wording was 0.94 on a 0 to 1 scale. The brand often still appeared, but the reason the AI gave for choosing it changed substantially.
Which modifier words caused the biggest shifts? +
Price words drove the largest rewrites, with 'cheapest' and 'most affordable' producing a mean word shift of about 0.97. 'Best' changed the underlying reasoning the most, with a median semantic distance of 0.60, while 'most trusted' produced the biggest positive swing in sentiment, moving it up by 0.50 across 16 observations.
Does the brand disappear from the answer, or just get repositioned? +
Usually it stays present but is repositioned. In one Home Insurance example, adding 'safest' turned AAMI from a brief 'large mainstream insurer' description into a detailed value and coverage rationale, whereas a simple mention count would have reported no change at all.
How is the positioning shift measured? +
Three metrics on a 0 to 1 scale: word shift, which captures how much of the visible wording changed using Jaccard distance; semantic distance, or meaning shift, which captures how far the underlying framing moved using embedding similarity; and net sentiment shift, which is movement toward positive wording minus movement toward negative wording.
What should marketing and SEO teams do about brand positioning in AI search? +
Measure the message, not just the mention. Build query neighbourhoods around the decision words your customers use, such as price, trust, safety, ease, and fit, give each brand a resilience profile by modifier, and focus AEO and GEO strategy on being chosen rather than merely mentioned. This is what brand consideration tracks.
Is the study's data and code available? +
Yes. The full code and dataset are public and reproducible in the Somantra AEO-research GitHub repository, so you can inspect the queries and regenerate every figure and table in the study.
About the author
Arun Prasad
Founder, Somantra
Arun Prasad is the founder of Somantra, an AI search visibility platform for brands, where he writes about answer engine optimisation (AEO) and AI search. His research analyses how brands surface in AI answers across ChatGPT and Google AI Overviews, including Somantra's studies of the Australian insurance market. He focuses on measuring brand visibility through systematic, large-scale conversational testing rather than one-off screenshots.
About the author
Yajat S
Yajat S is a researcher at Somantra, where he analyses how brands surface in AI search across ChatGPT and Google. He authored Somantra's Australian insurance citation study covering 2.4 million AI search citations, focused on the content formats and URL patterns that get cited by answer engines.