The list is the measurement
Change the questions and every figure changes with them. So the list has to be fixed before the first run and left alone until the recheck, otherwise two rounds cannot be compared. Treat it the way you would treat a survey instrument.
Where questions come from
Four sources, in order of reliability. Sales call recordings and pre-sales chat logs, because those are questions people actually asked. Search console queries that already bring traffic. Support tickets, which surface the doubts that block a purchase. Competitor comparison pages, which tell you what the market thinks the alternatives are.
What does not belong: keyword tool exports. A keyword is not a question. "GEO optimisation" tells you nothing about what someone wants to know, and a model asked that will produce a definition, not a recommendation.
Separate the brand-named questions
A question that already contains your brand name is not measuring visibility. Ask a model "is Brand X reliable" and it will say Brand X, because you put it there. Counted together with category questions, a handful of brand-named questions will lift the whole figure and the lift is entirely artificial.
Keep them, measure them, report them separately: they answer a different and genuinely useful question, which is how a model characterises you when someone already knows your name.
How many is enough
Enough that one question changing its mind does not move the headline number. Below roughly twenty, a single answer is worth five percentage points and the series will look volatile for no real reason. Beyond a hundred, the extra questions tend to be variations of ones you already have.
Freeze and version
Record the list with a version number and a date. When you add questions later, start a new version rather than extending the old one, and state in the report which version each round used. Comparing a 40-question round against a 60-question round is the most common way a measurement series quietly becomes meaningless.