Skip to content
Frank Vitetta

N° 002

Field note / Measurement

Prompt length vs brand mentions: 8,555 ChatGPT responses

A plateau, a step and a cliff, plus one bucket that depends on where the line is drawn. This note takes apart one small data set of my own. It covers the query shape, the table, a chart and the checks the table had to pass. It also explains why the result is a correlation and not yet a finding about length.

Author
Published
Last updated
Reading time
8 min

§ 1 The question

I wanted to know whether the length of a prompt goes with how often an AI assistant names a particular brand in its answer. My own tracking data holds 8,555 ChatGPT responses collected over nine months. They let me look.[1] This note takes that data set apart.

One metric runs through all of it. Each prompt has one tracked brand. That is the brand being monitored when the prompt ran. Mention rate is the fraction of responses that named that brand. A response that lists five other companies and omits the tracked one scores zero. A response that names it scores one, whether or not it speaks well of it.

The data covers a limited set of tracked brands and buyer-intent prompts. Everything below describes this data set only. It does not describe ChatGPT in general or other assistants.

Disclosure. This is my own data. So I have a commercial interest in the subject. Nobody else has reviewed the data set. It is not published. Details are on the method page.

§ 2 Method: bucketing by length

Each response carries the length of its prompt, counted in characters after leading and trailing whitespace is stripped. The shortest prompt in the data is 20 characters. The responses are grouped into six length buckets. Each bucket has a count of responses and a count of those that named the tracked brand.

A query that produces such a table has a simple shape. The snippet illustrates the approach with generic table and column names. It is not the production query.

IllustrationNot the production query

-- Illustration only: generic names, not the production query.
SELECT
  CASE
    WHEN char_length(trim(prompt_text)) < 35  THEN '1: 20 to 34'
    WHEN char_length(trim(prompt_text)) < 50  THEN '2: 35 to 49'
    WHEN char_length(trim(prompt_text)) < 70  THEN '3: 50 to 69'
    WHEN char_length(trim(prompt_text)) < 90  THEN '4: 70 to 89'
    WHEN char_length(trim(prompt_text)) < 120 THEN '5: 90 to 119'
    ELSE '6: 120 and over'
  END AS length_bucket,
  COUNT(*) AS responses,
  SUM(CASE WHEN brand_named THEN 1 ELSE 0 END) AS mentions,
  AVG(CASE WHEN brand_named THEN 1.0 ELSE 0.0 END) AS mention_rate
FROM responses
WHERE assistant = 'chatgpt'
GROUP BY 1
ORDER BY 1;

§ 3 The table and the chart

The result is six buckets, with the counts behind each rate.[1]

Table 1. Mention rate of the tracked brand by prompt length. ChatGPT, 8,555 responses collected over nine months
Prompt length (characters)ResponsesTracked brand namedMention rate95% intervalAgainst 50 to 69
20 to 341273930.7%23.4 to 39.234% lower
35 to 491,41962644.1%41.6 to 46.75% lower
50 to 693,0101,40546.7%44.9 to 48.5baseline
70 to 891,81264135.4%33.2 to 37.624% lower
90 to 1191,04029228.1%25.4 to 30.940% lower
120 and over1,147756.5%5.2 to 8.186% lower
All8,5553,07836.0%

The intervals are 95% Wilson intervals computed from the counts. The last column is relative. It sets each bucket's rate against the rate of the 50 to 69 bucket. Section 7 explains why the intervals are too narrow.

Plate 002.1Chart of Table 1

Mention rate of the tracked brand, by prompt length

ChatGPT / 8,555 responses / length in characters / n is the number of responses

Boundary-sensitive 20 to 34

  1. 20 to 34 n = 127

    30.7% 95% interval 23.4 to 39.2

Plateau 35 to 69

  1. 35 to 49 n = 1,419

    44.1% 95% interval 41.6 to 46.7

  2. 50 to 69 n = 3,010

    46.7% 95% interval 44.9 to 48.5

Step down 70 to 119

  1. 70 to 89 n = 1,812

    35.4% 95% interval 33.2 to 37.6

  2. 90 to 119 n = 1,040

    28.1% 95% interval 25.4 to 30.9

Cliff 120 and over

  1. 120 and over n = 1,147

    6.5% 95% interval 5.2 to 8.1

  • Bar: mention rate
  • Whisker: 95% Wilson interval
  • Dashed line: all 8,555 responses, 36.0%
  • Dashed outline: depends on the bucket boundary (section 5)
Figure 1 The figures from Table 1, drawn to one scale. The bars show the mention rate. The whiskers show the 95% Wilson interval. The dashed line marks the rate across all 8,555 responses. The 20 to 34 bar has a dashed outline because its value depends on where the bucket boundary is drawn. Section 5 shows this. The intervals assume independent responses, so the real uncertainty is wider than the whiskers.

§ 4 The check: does it reconcile?

Before reading anything into a table, I check that it agrees with a number it was not built from. There are two tests. The six response counts must add up to the 8,555 responses in the data set. The six mention counts must add up to the ChatGPT total in a separate company table that compares assistants. That total is 3,078 of 8,555, which is 36.0%.[1]

Table 2. Reconciliation: bucket counts against the independent ChatGPT total
Prompt length (characters)ResponsesTracked brand namedNamed ÷ responses
20 to 341273930.7%
35 to 491,41962644.1%
50 to 693,0101,40546.7%
70 to 891,81264135.4%
90 to 1191,04029228.1%
120 and over1,147756.5%
Sum of the six buckets8,5553,07836.0%
Cross-assistant table, ChatGPT8,5553,07836.0%

The sums and the right-hand column are my own arithmetic from the counts in Table 1.

Both tests pass exactly, with 8,555 responses and 3,078 mentions.

The check has already earned its place. An earlier draft of this study carried a different table, with different buckets and counts. Its rows did not reconcile with the 3,078 total, so it was thrown out. I do not reproduce it here.

Passing does not prove the table right. It shows that two tables count the same responses in the same way. The next section shows what a table can still hide after it passes.

§ 5 The boundary: one cut, two headlines

Table 1 has a weak shortest bucket. It shows 30.7% at 20 to 34 characters, against 44.1% beside it. Read quickly, that says very short prompts do worse. A second version of the same table says otherwise.

Before the database re-run that produced Table 1, the two shortest buckets were cut at 30 characters instead of 35. That version reconciles too.

Table 3. The same 1,546 responses, cut at 30 characters and at 35
BoundaryShortest bucketSecond bucketBoth together
Cut at 30Under 30: 39 of 87, 44.8% (interval 34.8 to 55.3)30 to 49: 626 of 1,459, 42.9%665 of 1,546
Cut at 3520 to 34: 39 of 127, 30.7% (interval 23.4 to 39.2)35 to 49: 626 of 1,419, 44.1%665 of 1,546

Each cell gives the responses that named the tracked brand, out of all responses in the bucket. The last column is my own addition.

The arithmetic leaves no room. Moving the boundary from 30 to 35 moves 40 responses into the shortest bucket (127 minus 87). It moves no mentions with them. 39 stays 39 and 626 stays 626. So the prompts of 30 to 34 characters produced 40 responses. Not one of them named the tracked brand. That figure comes from subtracting the two versions of the table. It is not a row printed in either.

Those 40 responses are the whole of the weak bucket. Put them with the shortest prompts and the bucket reads 30.7%. Leave them in the second bucket and the shortest prompts read 44.8%, level with the middle of the table. A five-character band of 40 responses may well be a handful of prompts run repeatedly. That is a hypothesis. The table does not say how many prompts are behind them.

So the careful reading is a narrow one. The table does not establish that short prompts do worse. It shows three things. One small band did worse. Prompts under 30 characters on their own sit at 44.8%. A bucket boundary can manufacture a headline or erase it.

There are three tables to keep apart. The draft in section 4 failed the check and was discarded. The cut at 30 reconciles and is the same data. The cut at 35 reconciles and comes from the database re-run. It is Table 1.

§ 6 The shape: plateau, step, cliff

Between 35 and 69 characters the rate is at its highest and close to flat. It is 44.1%, then 46.7%. Together those two buckets hold 2,031 mentions in 4,429 responses, which is 45.9%. They make up 51.8% of the data set (4,429 of 8,555).

From 70 characters the rate steps down. It is 35.4% at 70 to 89, then 28.1% at 90 to 119. At 120 characters and over it falls to 6.5%. That is about seven times lower than the 50 to 69 bucket (46.7 divided by 6.5 is 7.2).

So the shape is a plateau, a step and a cliff. The 20 to 34 bucket sits below the plateau at 30.7%. After section 5 I treat it as an open question. I do not treat it as the start of a curve.

§ 7 Why the intervals are optimistic

The Wilson calculation treats every response as an independent draw. These are not. They are repeated runs of a limited set of prompts and tracked brands. Answers to the same prompt are likely to resemble one another. That is clustering. Clustered data holds less information than its row count suggests. So the real uncertainty is wider than the intervals in Table 1 and the whiskers in Figure 1.

Nobody can work out how much wider from a summary table. The sound way is to resample whole prompts instead of single responses and recompute the rates each time.

Until that is done, my reading is this. The drop to 6.5% is too large for wider intervals to erase. The step at 70 characters is probably real in this data set. The gap between the two plateau buckets should be ignored because their intervals already overlap. The 20 to 34 bucket stays unresolved.

§ 8 The confound: length or wording?

This is a correlation. The study did not control for what the prompts ask. Length is tangled up with intent and wording. The study's own description says the long prompts in this data set tend to be narrative, advice-seeking questions with several constraints. The short ones tend to be list requests of the "best X" kind.

So the table cannot say that length causes the drop. One plausible explanation is that each extra constraint narrows the set of brands that fit. Another is that a long, personal question draws advice in reply instead of a list of names. Both are hypotheses. Neither has been tested.

Three more limits apply. Character count is a crude ruler. A word count or token count may be a better one. The table covers ChatGPT only. It describes a limited set of tracked brands and buyer-intent prompts, so a different set could give a different curve.

§ 9 The experiment that would settle it

Two designs would separate length from wording.

  1. Hold the question fixed and vary only the length. Write each question at several lengths without adding or removing a constraint. Run every version the same number of times. Compare within each question.
  2. Model length and wording type together. Label each prompt by type (list request, comparison, advice). Then fit one model of mention against both length and type. Group responses by prompt so that repeated runs are not counted as independent.

In either design I would measure length in tokens as well as characters. I would treat it as a continuous value so that no boundary has to be chosen. I would repeat the test on other assistants.

§ 10 What to take from it today

One point survives every caveat. If you build test prompt sets to measure how assistants answer, the length mix of the set moves the result. The overall 36.0% in this data set is a blend of buckets that run from 46.7% down to 6.5%. Leave out the 120-and-over bucket and the rest is 3,003 mentions in 7,408 responses, which is 40.5%. Nothing about the assistant has changed. Only the mix has.

In practice:

  • Record prompt length as a column in every test set, in characters and in tokens.
  • Keep the length mix fixed between runs. Otherwise report each length bucket alongside the overall rate.
  • State your bucket boundaries. Cut the table a second way before trusting a small bucket.
  • Reconcile every bucketed table against an independent total before anyone else sees it.

I would not conclude that shorter prompts earn mentions. I would not conclude that the shortest ones lose them either. This table does not show that length is the cause of either.

Sources

The figures in this note were checked against the study tables on 5 October 2026.

  1. My own internal study tables on prompt length and tracked-brand mentions in ChatGPT responses. They are the bucket table from the database re-run, the earlier version of it cut at 30 characters and the cross-assistant summary table used for the reconciliation check. Data collected over nine months. The underlying data set is not public and has not been independently reviewed. There is no published version of the study to link to.