The differences that actually matter

Three, in order of practical impact. Whether the platform searched the web for this particular answer, which varies by question and sometimes by run. Whether it attributes sources in the answer body or not at all. How long the answer is, because a short answer names fewer brands and position within it matters more.

Record search state with every run

An answer composed from training data and one composed from a live search are different measurements. Mixing them produces a series where improvements appear and vanish for reasons that have nothing to do with your work. Record, per run, whether search was used - and if the platform does not expose that, record the setting you used.

Citations: count only what the answer attributes

Some platforms expose a list of pages they consulted alongside the answer. That list is a retrieval candidate pool, not a set of citations - most of it was not used. Counting the pool as citations inflates the figure several times over and makes a bad source look influential. Count only sources the answer body actually attributes.

What does not differ

The content work. A page with specifics, reachable by crawlers, corroborated elsewhere, performs better everywhere. There is no platform for which thin content works, and no per-platform trick that substitutes for the fact base.

Where the difference does change decisions

Placement. Platforms weight sources differently, and an outlet that one platform retrieves constantly may be invisible to another. If your measurement shows a persistent gap on one platform, look at which sources that platform does cite in your category and place there, rather than writing more of the same.