Field Notes
Set the Question Before the Answer Arrives
Suppose we stage an exhibition to test whether 100 people will buy $15 tickets.
Thirty-eight tickets sell.
That is disappointing.
Then we notice that 140 people visited the event page.
Social media engagement was strong.
The artists enjoyed the experience.
Several people told us the exhibition looked beautiful.
All of those things may be true.
But none changes the result of the question we originally asked.
Do not change the question because you dislike the answer.
If the Cultural Interest Study is going to produce defensible findings, we need to know what we are trying to learn before the result arrives.
Write down the question first
Before an activation opens, we should be able to finish a simple sentence:
We are doing this to learn whether…
The ending will depend on the test.
…customers will purchase 75 tickets at $25.
…a free introductory experience produces later paid purchases.
…buyers who attended the first program return for the second.
…a different location produces a different purchasing pattern.
…customers continue buying when the price increases.
…a sponsor supports the program again.
Those are different questions.
They require different evidence.
Write down the question before opening the Box Office.
Evaluation begins before the result
This is standard evaluation discipline.
The Centers for Disease Control and Prevention’s 2024 Program Evaluation Framework places the evaluation questions and design before evidence collection, conclusions, and action.
The framework asks evaluators to identify the purpose of the evaluation, the questions it should answer, who will use the findings, and how those findings are expected to be used.
Only then does the framework move into gathering evidence and deciding which measures will answer those questions.
The question gives the evidence a job.
Without that job, we can collect enormous amounts of information without knowing which information matters.
The measure should answer the question
If the question is about paid demand, likes on Instagram do not answer it.
Purchases do.
If the question is about attendance, tickets sold do not completely answer it.
People through the door do.
If the question is about repeat demand, one successful event does not answer it.
Another purchase does.
If the question is about sponsor renewal, the first sponsorship does not answer it.
The sponsor supporting something again does.
The measure should describe the behavior the question is actually asking about.
Questions and measures are not the same thing
The question can be broad.
The measure makes it observable.
Question
Will customers pay for this exhibition?
Primary measure
Number of paid tickets purchased at the tested price.
Or:
Question
Do first-time customers come back?
Primary measure
Percentage of first-time buyers who make another purchase during the defined follow-up period.
Or:
Question
Does a free experience create later paid demand?
Primary measure
Number or percentage of identifiable free-event participants who later complete a paid purchase during the defined period.
The Cultural Interest Study should establish that connection before collecting the result.
We can measure many things without making them all primary
An activation can generate a large amount of information.
Revenue.
Attendance.
Show rate.
Website traffic.
Email response.
Social engagement.
Customer location.
Sponsor activity.
Nearby spending.
Repeat purchases.
All of that may be useful.
But not every measure should get to redefine the purpose of the test.
The CDC framework warns that collecting too many indicators can distract from the goals of an evaluation and consume resources without improving the answer.
Measure broadly when useful. Decide in advance which measure answers the primary question.
Set an expectation
The next step is harder.
Before we know the result, what result would matter?
The CDC framework recommends establishing expectations for the type and level of evidence that will be used to answer an evaluation question.
That gives the eventual result a point of reference.
Suppose a room holds 100 people.
We are testing a $25 ticket.
Based on production cost, prior experience, capacity, and the decision we are trying to make, we might establish an expectation before sales begin.
Perhaps we believe:
Fewer than 40 purchases would justify substantial revision.
Between 40 and 70 would justify another test.
More than 70 would justify repeating or deepening the offer.
Those numbers are only an example.
The actual expectation should have a reason behind it.
A threshold needs a reason
We should not invent a number because it sounds scientific.
Why would 70 tickets matter?
Perhaps:
Seventy tickets cover a meaningful share of production cost.
Seventy represents a strong comparison with a previous test.
Seventy demonstrates enough demand to justify a larger room.
Seventy is the minimum needed for a particular operating model.
If crossing the threshold does not change anything we would do, the threshold may not be useful.
A threshold matters when crossing it changes a decision.
Do not turn every test into pass or fail
The Cultural Interest Study does not need to reduce every result to:
Success.
Failure.
Markets rarely behave that neatly.
Suppose our expectation was 75 paid tickets and we sell 72.
It would be absurd to declare that three additional tickets would have transformed a failed market into a successful one.
Ranges can often be more useful than rigid pass-or-fail lines.
The expectation should help us interpret and act, not manufacture certainty the evidence does not contain.
Ask what the result would make us do
Evaluation is most useful when the answer can inform a decision.
The CDC framework emphasizes use throughout the evaluation process and ends with acting on findings.
That suggests another question to ask before an activation:
What result would change what we do next?
Perhaps:
Strong demand means repeat the test at a larger capacity.
Moderate demand means test another price.
Weak demand means revise the offer substantially.
Strong free attendance but weak conversion means reconsider the path from free participation to paid purchase.
Strong first purchase but weak return means investigate repeat behavior.
Now the test has a purpose beyond producing a number.
Some results will change the question
This does not mean the study can never learn something unexpected.
It should.
Suppose an exhibition performs weakly at the Box Office.
But nearly half the buyers later purchase another Foundation program.
That was not the primary question.
It may still be an important observation.
It can generate the next question.
An unexpected finding can create a new test. It should not rewrite the test that already happened.
Primary, secondary, and exploratory findings
We can make that distinction explicit.
Primary finding
What happened to the measure selected to answer the main question?
Secondary finding
What happened to other measures we intentionally chose to collect?
Exploratory observation
What unexpected pattern appeared that may deserve a future test?
All three can be useful.
They are not interchangeable.
The unexpected result gets to inspire the next question. It does not get to replace the first one.
Success theater can weaken the evidence
Arts organizations have many reasons to want an event to sound successful.
Artists worked hard.
Funders supported it.
Staff spent time on it.
Volunteers helped.
Nobody enjoys announcing that a test produced a weak result.
That creates a temptation to find another measure.
Social media performed well.
The room looked good.
Someone important attended.
We received compliments.
The artists enjoyed themselves.
Those observations can have value.
They cannot substitute for the primary measure if the primary question was whether people would buy.
A meaningful experience and a successful market test are not automatically the same finding.
Return to the exhibition
We began with an exhibition.
The question was whether 100 people would buy $15 tickets.
Thirty-eight purchased.
The primary result is:
Thirty-eight customers purchased the exhibition at the tested $15 admission price.
Then we can add:
Website traffic was strong.
Social engagement was high.
Perhaps 34 of the 38 buyers actually attended.
Perhaps several purchased something else later.
Those are additional findings.
The report becomes richer without changing the original answer.
Strong results need discipline, too
The same rule applies when we love the answer.
Suppose we expect 75 buyers.
We get 160.
Excellent.
The result substantially exceeded the expectation.
That deserves attention.
But we still need to remember what we tested.
Maybe the artist brought a large following.
Maybe tickets were priced too low.
Maybe the location contributed.
Maybe the marketing reached an unusually strong audience.
The correct response to an unexpectedly strong result may be another test.
A result that exceeds expectations is evidence to deepen, not permission to stop asking questions.
Set the objective before designing the test
The National Institute of Standards and Technology makes the same point in formal experimental design.
Planning begins by identifying the objective.
NIST recommends writing down the objectives, prioritizing them, then selecting the factors, responses, and design needed to answer them.
The Cultural Interest Study is not a laboratory experiment.
But the discipline transfers well.
First decide what you need to know. Then decide what you need to observe.
The question should be narrow enough to answer
Consider:
Do Duncanville residents support the arts?
That question is so broad that almost any evidence could be made to answer it.
Compare it with:
Will at least 60 customers purchase tickets to this 90-minute jazz performance at $25 under the documented conditions?
Now we know what evidence matters.
The narrower question does not answer everything about jazz.
It answers one thing clearly.
Repeated clear answers can eventually support larger findings.
Do not ask a test to answer more than it can
A well-defined question also limits the conclusion.
Suppose 85 people purchase the $25 jazz performance.
We can say the test exceeded the expectation of 60 purchases.
We can describe the customers.
We can compare the result with similar tests.
We still should not leap immediately to:
Duncanville has a permanent market for jazz.
The question we asked was smaller than that.
So the answer should be smaller, too.
Pre-commitment protects objectivity
The CDC framework identifies independence and objectivity as an evaluation standard.
It also emphasizes transparency so that evaluation methods are not tailored to produce preferred findings.
That principle matters here.
If we decide which result counts only after we see the numbers, our preferences can influence the interpretation.
Setting the primary question, measure, and expectation in advance creates a record of what we meant to learn before we knew whether the answer would please us.
The point of setting expectations in advance is not to predict the answer. It is to prevent the answer from rewriting the question.
The activation record now has a before and after
Field Note 18 proposed a test record for every activation.
We can now make that record more precise.
Before launch:
Primary question
What are we trying to learn?
Primary measure
What observable behavior will answer that question?
Expectation
What range of results do we reasonably anticipate or consider meaningful?
Decision rule
What kinds of results would make us repeat, revise, expand, or stop?
Conditions
What exactly are we offering, to whom, at what price, place, time, capacity, and marketing level?
Afterward:
Primary result
What happened to the measure?
Difference from expectation
How did the observed result compare with what we established beforehand?
Secondary findings
What else did our planned measures show?
Exploratory observations
What unexpected patterns appeared?
Limitations
What prevents us from making a larger claim?
Next action
What should we repeat, revise, expand, investigate, or stop?
That record turns an activation into a test whose meaning does not depend on our mood afterward.
We can revise the plan when reality requires it
Setting the question beforehand does not make the study rigid.
The CDC framework itself recognizes that evaluation planning can change when feasibility, available evidence, or context changes.
Suppose a venue closes unexpectedly.
We relocate the program.
That changes the test.
We should document the change.
Suppose severe weather changes attendance conditions.
Record it.
Suppose the artist withdraws and another replaces them.
Record it.
Transparency matters more than pretending the original plan remained untouched.
A living protocol can change. It should leave a record of what changed and why.
Do not bury the weak result
There will eventually be tests that perform poorly.
Those findings belong in the study.
If a planned paid-demand test produces weak purchasing, report it.
If a free event produces a high no-show rate, report it.
If a sponsor does not renew, record it.
If an audience does not return, record that, too.
Those results may be precisely what protects the Foundation, artists, businesses, property owners, and future investors from making expensive decisions based on enthusiasm rather than evidence.
A study that can report only good news is not much of a study.
The Board should be able to see the original question
This also improves accountability.
When the Foundation reviews an activation later, the Board should not have to rely on our memory of what we hoped to accomplish.
The record should show:
What we measured.
What we expected.
What happened.
What we concluded.
What we did next.
That creates a chain from intention to evidence to decision.
The public finding can remain simple
None of this means every public report has to read like a research journal.
The internal record can carry the detail.
The public finding can still be clear.
“The Foundation tested a $25 ticket for this jazz performance with an expectation of at least 60 purchases. Eighty-five tickets were purchased. The result exceeded the pre-established expectation. A second jazz offer will test whether similar purchasing repeats with another artist.”
That is understandable.
It is also defensible.
The reader knows what we tested, what happened, and why we are doing something next.
Three questions now protect the study
The last three Field Notes give us a methodological spine.
First:
Who does this evidence actually represent?
Second:
What exactly did we test?
Third:
What question were we trying to answer before we saw the result?
If we can answer all three, our finding has boundaries.
We know who it describes.
We know what they responded to.
We know what behavior mattered before we learned whether we liked the outcome.
Good evidence begins before the evidence arrives.
Ask the question.
Name the evidence.
Set the expectation.
Run the test.
Report what happened.
Then decide what comes next.
Do not decide what the test meant only after you discover whether you liked the result.
Subscribe to Ron Thompson’s Field Notes
Occasional notes on art, culture, creative economies, and the work of building cultural capacity from the ground up.
Subscribe to Field NotesBy sending the subscription request, you agree to receive Ron Thompson’s Field Notes by email. Your email address will be used to manage and deliver your subscription and will not be sold. You may unsubscribe at any time by replying “unsubscribe.”
Sources
Centers for Disease Control and Prevention. “CDC Program Evaluation Framework, 2024.” Morbidity and Mortality Weekly Report, vol. 73, no. RR-6, 26 Sept. 2024.
Centers for Disease Control and Prevention. “Evaluating Your Programs.” Health Literacy, 16 Oct. 2024.
National Institute of Standards and Technology. “What Are the Objectives?” NIST/SEMATECH e-Handbook of Statistical Methods.
National Institute of Standards and Technology. “What Is Experimental Design?” NIST/SEMATECH e-Handbook of Statistical Methods.
National Institute of Standards and Technology. “What Are the Steps of DOE?” NIST/SEMATECH e-Handbook of Statistical Methods.

