Field Notes
What Exactly Did We Test?
Suppose the Duncanville Arts Foundation presents a contemporary dance performance.
Tickets are $35.
It happens Thursday at 7:30 p.m.
The performer is unfamiliar to most local buyers.
Marketing begins ten days before the event.
Forty tickets sell.
What did we learn?
It would be easy to say:
There is limited demand for dance in Duncanville.
But that statement is larger than the test.
What we actually observed was forty purchases for one particular dance offer under one particular set of conditions.
The market does not respond to a category in the abstract. It responds to an offer.
The offer is the test
We say we are testing theater.
Or jazz.
Or film.
Or exhibitions.
Those categories are useful for organizing the Cultural Interest Study.
But the customer never encounters “theater” as an abstract research category.
The customer encounters something much more specific.
A particular production.
A particular artist.
At a particular price.
In a particular place.
At a particular time.
Presented in a particular way.
That bundle is what the customer accepts or rejects.
Every activation contains several conditions
The Cultural Interest Study should document those conditions.
Product
What exactly was presented?
Artist
Who created or performed it?
Price
What did admission cost?
Place
Where did it happen?
Time
What day and hour?
Format
How was the experience structured?
Marketing
Who had a reasonable opportunity to hear about it?
Capacity
How many purchases were even possible?
Those conditions help define what we tested.
A ticket sale is a response to the bundle
Suppose someone pays $35 for the dance performance.
What exactly did they vote for with that purchase?
Dance?
That artist?
That venue?
A Thursday-night activity?
A friend who invited them?
The answer may be several things at once.
The transaction proves that the complete offer was acceptable enough for that customer to buy.
It does not automatically tell us which part of the offer caused the purchase.
A ticket sale is a vote for the offer that existed, not every possible version of the idea behind it.
The same is true when someone says no
A weak sale is also a response to the bundle.
Suppose only forty people buy.
Maybe the art form had limited appeal.
Maybe $35 was too high.
Maybe Thursday was inconvenient.
Maybe ten days was not enough time to market an unfamiliar artist.
Maybe the location created friction.
Maybe several of those things mattered.
One result does not automatically identify the cause.
Failure is a result. The conditions help us investigate the explanation.
Compare two jazz tests
Imagine two concerts.
Test A
Tuesday night.
$40 ticket.
Unfamiliar venue.
35 buyers.
Test B
Saturday night.
$25 ticket.
Familiar restaurant.
110 buyers.
Which one represents the Duncanville jazz market?
Neither by itself.
Together, they tell us that two jazz offers produced very different results under different conditions.
That difference gives us something to investigate.
Price may matter.
Place may matter.
Day may matter.
The performers may matter.
Several factors may interact.
Conditions can change the response without changing the category.
Real-world tests are not laboratory experiments
The Cultural Interest Study operates in real places with real artists and real customers.
We cannot freeze every condition.
Weather changes.
Artists change.
People have schedules.
Locations differ.
Marketing never reaches exactly the same people twice.
That does not prevent us from learning.
It means we need to document context and be careful about causation.
The Centers for Disease Control and Prevention’s 2024 Program Evaluation Framework makes context the first step of evaluation and notes that location, environment, people, and other surrounding factors can affect both outcomes and how findings should be interpreted.
Context is not noise to be ignored. It is part of understanding what happened.
Experiments begin with factors and responses
The National Institute of Standards and Technology describes experimental design in a simple way.
An experiment deliberately changes one or more factors and observes what happens to one or more responses.
In the Cultural Interest Study, possible factors include:
Artist.
Art form.
Location.
Day.
Start time.
Format.
Promotion.
Possible responses include:
Revenue.
Reservations.
Attendance.
Show rate.
Additional spending.
Repeat purchase.
That vocabulary is useful even when our field tests are much simpler than formal laboratory experiments.
Some factors are outside our control
NIST also recognizes that experiments can contain uncontrolled factors.
Cultural activity certainly does.
Rain begins an hour before an outdoor concert.
A major sporting event happens the same night.
Road construction makes a venue harder to reach.
A performer suddenly receives regional media attention.
None of those conditions may have been part of the planned test.
But if we know they occurred, we should record them.
An unexpected condition does not erase the result. It becomes part of the record we use to interpret it.
Change fewer things when the question is narrow
Suppose the specific question is:
Does this audience respond differently at $20 than at $30?
A useful comparison becomes harder if we also change the artist, venue, day, format, and marketing strategy.
If nearly everything changes, a different result may be impossible to attribute to price.
Real-world testing will not always allow clean comparisons.
But when we are trying to understand one particular factor, keeping other important conditions reasonably comparable can make the result easier to interpret.
The narrower the question, the more useful a disciplined comparison can become.
Changing several factors can still be useful
There are also times when we intentionally want to change the entire offer.
Perhaps the first version failed.
We redesign it.
New place.
New price.
New format.
New marketing.
If the second version performs much better, that is useful.
We learned that the redesigned offer produced a different result.
What we may not know is which individual change caused the improvement.
We can learn that a bundle worked better without pretending we isolated the reason.
A sold-out event can mislead us, too
This discipline is not only for weak results.
Imagine a folk music performance sells out instantly.
That appears to show strong demand for folk music.
Then we discover the performer has 20,000 regional followers and personally promoted the event.
What did we test?
We certainly observed strong demand for that offer.
But perhaps much of the demand belonged to that particular artist.
The next useful test might be another folk performer.
Similar price.
Comparable capacity.
Similar setting.
Now we have another observation.
Success deserves the same scrutiny as failure.
Capacity can create false impressions
Suppose a room holds 40 people.
Forty tickets sell.
Sold out.
That proves at least 40 buyers accepted the offer.
It does not tell us whether demand stopped at 40.
Perhaps another 100 people would have purchased.
If the Box Office closes sales when capacity is reached, those additional transactions cannot occur.
A waitlist can give us some additional information.
Capacity can place a ceiling on what the transaction data are able to reveal.
Marketing reach defines part of the test
Consider another example.
We sell only 30 tickets.
But our promotion reached only 150 people.
That is a different situation from selling 30 tickets after 10,000 relevant people repeatedly encountered the offer.
The final sales number is the same.
The exposure behind it is not.
This is why the study should record what we can about how an offer was promoted.
We may not be able to know exactly who saw every message.
But we can document:
When promotion began.
Which channels were used.
Whether paid advertising was used.
How much was spent.
Whether partners promoted the offer.
Any available reach or response data.
The customer also has alternatives
An arts offer does not exist in isolation.
The U.S. Small Business Administration recommends looking at demand, location, pricing, market size, and alternative options when studying a market.
That is relevant here.
A customer deciding whether to buy our $30 concert ticket may also be deciding between:
A movie.
Dinner.
A sporting event.
Staying home.
Spending the money on something unrelated to the arts.
The offer competes for both money and time.
The conditions surrounding the test therefore include more than what happens inside the venue.
One test should not define an art form
This should become a formal discipline of the Cultural Interest Study.
One weak theater offer should not become:
Duncanville does not support theater.
One sold-out jazz event should not automatically become:
Duncanville has a large jazz market.
One successful exhibition should not become:
Residents prefer visual art.
Those may eventually become defensible findings in some form.
One activation rarely gets us there.
Repeat the category with a different offer
If we want to learn about the broader category, we need more than one version of it.
Suppose we test theater several times.
A comedy at $25.
A drama at $30.
A short-form performance at $15.
A staged reading at $10.
Different audiences may appear.
Different price sensitivities may appear.
Perhaps one form performs consistently better.
Perhaps all of them struggle.
Now our understanding of the category begins to rest on more than a single product.
A market becomes clearer when more than one offer points in the same direction.
Describe the result before explaining it
The reporting sequence matters.
First:
What happened?
Then:
What might explain it?
Then:
What should we test next?
Keeping those steps separate helps prevent interpretation from becoming fact.
Use language that matches the test
Consider the difference.
Too broad:
“Residents will not pay $30 for theater.”
More precise:
“This theater offer generated 42 purchases at a $30 ticket price.”
Or:
Too broad:
“There is strong demand for live music.”
More precise:
“Three live music tests produced repeated purchasing at the tested prices, locations, and capacities.”
Or:
Too broad:
“People do not want dance.”
More precise:
“This dance test did not generate sufficient purchasing under the conditions tested.”
Name the result before you name the explanation.
Every activation should have a test record
The Cultural Interest Study should create a simple record for each activation.
Before the event:
What are we offering?
What question are we trying to answer?
What is the ticket price?
What is the capacity?
Where and when will it happen?
How will it be marketed?
What outcomes will we measure?
After the event:
How many people purchased?
How much revenue was generated?
How many attended?
What was the show rate?
What did the activation cost?
What surrounding conditions may have mattered?
What can we reasonably conclude?
What should we test next?
Document what changed
When we repeat an art form or revisit an offer, the study record should also identify what changed.
Price?
Artist?
Location?
Marketing?
Day?
Time?
Capacity?
Format?
Then we can compare the results with more discipline.
The test record turns an event into evidence we can return to later.
Some conclusions will remain uncertain
This is important.
Sometimes we will not know why a result occurred.
We may have several plausible explanations.
The evidence may not separate them.
That is acceptable.
A defensible study can say:
We observed this result, but the available evidence does not establish which condition caused it.
That is much stronger than selecting an explanation because it sounds reasonable.
Field research should stay kinetic
The purpose of this discipline is not to make the Cultural Interest Study rigid.
It should do the opposite.
An observation creates another question.
The question creates another test.
The next result either strengthens the pattern or complicates it.
Then we adjust again.
Audience-building research supported by the Wallace Foundation has used a similar cycle of research, implementation, assessment, and revision.
The point is to learn while doing.
The study should become more precise as it accumulates experience, not more certain than the evidence allows.
Two questions should follow every finding
Field Note 17 established one question:
Who does this evidence actually represent?
Now we add another:
What exactly were those people responding to?
Together, those questions place boundaries around the finding.
Who did we observe?
What did we offer them?
Under what conditions?
What did they do?
What can we reasonably say because of it?
Every finding should identify both the people observed and the conditions tested.
We did not test jazz in the abstract.
We did not test theater in the abstract.
We did not test exhibitions in the abstract.
We tested particular offers made to particular people under particular conditions.
Name the test before you name the market.
Subscribe to Ron Thompson’s Field Notes
Occasional notes on art, culture, creative economies, and the work of building cultural capacity from the ground up.
Subscribe to Field NotesBy sending the subscription request, you agree to receive Ron Thompson’s Field Notes by email. Your email address will be used to manage and deliver your subscription and will not be sold. You may unsubscribe at any time by replying “unsubscribe.”
Sources
National Institute of Standards and Technology. “What Is Experimental Design?” NIST/SEMATECH e-Handbook of Statistical Methods.
National Institute of Standards and Technology. “Overview: What Is Experimental Design?” NIST/SEMATECH e-Handbook of Statistical Methods.
Centers for Disease Control and Prevention. CDC Program Evaluation Framework, 2024. Morbidity and Mortality Weekly Report, vol. 73, no. RR-6, 26 Sept. 2024.
U.S. Small Business Administration. “Market Research and Competitive Analysis.” Business Guide.
Ostrower, Francie. In Search of the Magic Bullet: Results from the Building Audiences for Sustainability Initiative. The University of Texas at Austin and The Wallace Foundation, 2024.

