Showing posts with label Reliability Sample. Show all posts
Showing posts with label Reliability Sample. Show all posts

Thursday, July 10, 2014

Nielsen’s House Of Cards

A very cool aspect of the business Mike O’Malley, Becky Brenner and I are in is fully understanding and - many times - being a part of moving ratings. 

We make it our business to know what causes, just for a few examples, WMZQ/Washington to grow from a 2.7 a few months ago to a 4.1 6+ in the current month;  how WYCD/Detroit got back up to a 6.2 after a 5.3-5.5 trend and the driver of KNIX/Phoenix’s largest PPM share in the station’s history this month, a 6.2, after a 5.0 last month.

As a result, we were all extremely relieved to see Nielsen’s announcement yesterday on the Los Angeles sample that in spite of one media related household that resulted in Univision firing a manager at LA 102.9 “our analysis revealed that any significant differences in the estimates over the past year were isolated primarily to a single station” and didn’t impact much else. 

It turns out that the wobbly numbers affecting many other stations were a result of a UPS driver who listened to radio all day in his delivery truck and the rest of his family traveling to Mexico but leaving their meters at home in hopes they’d still be paid their premiums to help underwrite their vacation.

Two homes is all it took!

Lets face it.  Due to the very small sample sizes, especially when it comes to ultra core radio users, it’s as much an art as it is a science to keep PPM panels consistent and believable.

If Nielsen had been forced to rebuild their entire Los Angeles panel, it would have been expensive and a lengthy process, affecting immense revenues in the nation’s second largest metro.

Wednesday, May 14, 2014

Making "Bad" 6% Less So

Not to seem ungrateful that we're now just a few more months away from seeing the 6% increase in PPM samples which Nielsen promised for 2014 last November, but - as my post "The Canary Has A Name" suggests - that won't come even close to realizing the dream of true PPM-based cross media measurement with reliable estimates a reality.

As Paragon Research's Larry Johnson recalls "when Nielsen and Arbitron were first collaborating in the early 1990s about rolling out radio PPM jointly, the sample size was to be roughly three times what Arbitron has been offering.  The sample size becomes even more crucial if Nielsen is going to measure multiple platforms with a PPM device.  If you’re going to accurately measure outlets with small, niche audiences, the sample size must increase dramatically.  For example, cable television gets cut up into very small pieces given the huge number of cable channels available.  It’s interesting that Nielsen uses a mixed methodology utilizing both diaries and meters in many of the television markets they rate."

Researcher Mark Ramsey offers a brief statistics course in a two year old article to underscore why it's important that Nielsen make the most of their abilities now that that they own PPM: 

Small samples yield extreme results more than large samples do. Every broadcaster knows this, of course, but the impact is far worse than you think.

Daniel Kahneman discusses this issue in some depth in his book Thinking, Fast and Slow:

Imagine a large urn filled with marbles. Half the marbles are red, half are white. Jack draws 4 marbles on each trial, Jill draws 7. They both record each time they observe a homogeneous sample—all white or all red. If they go on long enough, Jack will observe such extreme outcomes more often than Jill—by a factor of 8.

By a factor of 8!  Extreme (and erroneous) results are 8 times more likely with 4 rather than 7.
And how many meters are tuned to your station right now?

Ramsey asks:   "What does it say about the veracity of our message and our messengers when we’re using numbers we know to be flat-out false?"

Another researcher, Kurt Hanson recalls his attempts several decades ago to sell radio on a better sample size in a service called Accu-Ratings, and puts it this way: "introduction of Portable People Meters (PPMs) solved the wrong problem:  Diaries weren’t being filled out precisely to the minute (in fact, lots of respondents filled out their diaries in one sitting at the end of the week), and PPMs “fixed” this problem, but because meters were so much more expensive than paper diaries (hundreds of dollars vs. maybe a dime), Arbitron used far fewer of them per market.  So the sample size issue got *worse,* not better!"

If today's multi-media ratings firms are serious about credibly measuring the increasingly-fragmented pie, they have to see the potential of signing many, many new clients which hopefully will rapidly grow their revenues, while keeping radio and television's rates stable at current levels.

I hope they also understand that to achieve that lofty goal sample sizes must be at least tripled or quadrupled, not merely increased in small increments.

Sunday, March 03, 2013

When Nielsen Takes Control Of Arbitron (Some Wishful Thinking)

During this legally-mandated silent period as Nielsen does what it takes to close on Arbitron later this year, no one is talking about what will actually happen to radio and TV ratings.

My hopes:
  • Nielsen solves its "set top box in a four screen mobile world" problem by moving to the Portable People Meter (either its own or ARB's, which ever can be accredited in every measured market) to get out of "household" TV ratings and move to individual measurement.
  • Arbitron stops hitting its sample targets by trying to find homes where they can place (many!!) more than 2-3 meters by going to one meter per household.
  • The sample for both will include cell phone only persons in their exact proportion to the general population.
  • The sample sizes are quadrupled (or even more) as the same meters follow all media use of one large, perfectly representative group of people at home, at work, in their vehicles, no matter where or how they access them.
  • Reliability of all estimates is doubled, breathing new life into low cume, high loyalty niche formats and decreasing the silly, unbelievable share compression infecting today's ARB numbers.
  • Stations with poor signals are rejuvenated, once again becoming viable businesses instead of low rate bottom-fishers.
  • The antique "radio diary" is retired forever and all 210 TV DMA's are measured by portable personal meters, tracking usage of all media in real time in all of them.
  • New media measurement opens the service to working for many emerging media clients, dramatically increasing their revenues.
  • Nielsen's Radio and TV rates are lowered.  Using the same panel for all media is more efficient.
  • National and network buys start to increase for the most-used radio stations in all markets, not just the major markets as is increasingly happening now based on this deeper understanding of radio's amazing reach everywhere.
  • Cross-media measurement proves that radio is 15-20% of the average person's media day, and at first media buyers attempt to keep cost per point levels and radio's share of media buys the same as they always have been, but ultimately radio revenues almost double over time as buyers can't deny the reach, usage and effectiveness of radio.
  • Nielsen saves money as they do all of this, making it possible for them to easily retire the debt they have incurred to do the deal.
  • Arbitron and Nielsen - as well as all of their existing clients - live happily ever after in a lovely marriage.
I can dream, can't I?  And, sometimes, dreams even come true.  Let's hope they do in this case.

Monday, October 22, 2012

What Every Music Director Knows That Arbitron Seemingly Doesn't

Big kudos to Lincoln Financial Media's Don Benson for his service in the past year as chair of the Arbitron Affiliates Advisory Council.  No one could have done more to call attention to the constant elephant in the room in radio's relationship with ARB, sample sizes.

The challenge appears to be that ARB would like the AAAC to advocate to their radio brethren that the only way to get sample sizes up is to increase the rates they pay and of course meanwhile radio tells their reps on the council that given that the ratings giant's profits have been growing faster than the economy even during these recessionary times and are locked in at the same levels for the next few years, so they see ARB as a parasite that is slowly killing its host.

Once upon a time, when radio research companies were these distant experts who also charged humongous fees for perceptual studies and music testing because they were the masters of techniques that were intellectually above the heads of mere mortals.

Today, that's no longer true, as online music testing and perceptual studies are routinely done in house using email databases and anyone who sees the results week after week can quickly observe the difference a large, representative sample creates in the results versus a too-small one.


This ranker, shared with permission, from a recent test tells me that a total sample of 126 is a pretty reliable indication of how the average listener to this station feels about these songs, especially so if the station looks at a weekly trend report from different people of about the same size sample over multiple weeks and the stats remain fairly stable.


Another station whose data I watch each week just had less than half the sample size of the first one.  Hopefully, the database manager/music director of this station knows that those columns with so many 100%'s - which look a lot like PPM rankers in some narrow cells - are not worth making decisions on.

I'd prefer to see a sample of 250 or more, so all of the narrow demo tabs are also well-balanced and at least the minimum of 30 persons that ARB's PD Advantage and Maximiser will allow you to run a report on. 

49 people with just five or six in the narrow cells? = garbage data. You'd be better of with no research than acting on this particular music test.

Yet, media buyers increasingly are looking at very granular cume and TSE numbers for individual radio stations based on samples that small and expecting them to be consistent and stable.

Hopefully, 2013 council Chair Craig Jacobus and Vice Chair Joel Oxley will be just as vociferous as Benson and his committee has been in 2012 and - for the good of all of us - finally move Arbitron toward not just listening to radio's long-stated concerns, but ACTION on out-dated paper-based diary methodology and sample size increases in all size markets.

See:  How Do You Keep Online Music Testing Sample Sizes Robust? (pdf) A&O music specialist Mark Patric shares his approach.  What's yours?

Wednesday, June 27, 2012

How Big Do Samples Need To Be?

The best way to be sure that a survey sample is reliable:  do repeated measurements consistently produce the same results?

For example, in Arbitron market #67 Fresno - where thankfully both owners still subscribe to ARB so we can at least get some idea from the real world - the top country station (Peak's KSKS) has gone from #3 in 12+ monthly trend cume rank in May 2011 to #5 to #2 to #3 to #3 to #8 to #7 to #4 to #3 to #5 to #6 to #8 to #3 in May 2012.

The second tier country station (Clear Channel's KHGE) wobbled even more over the same period, from #21 to #11 to #7, then #16, #14, #21, #16, #10, #15, #18, #12, #17, #9.

I chose 12+ cume rankers to trend, since those should be the largest, most-stable numbers to track, less susceptible to weighting-driven TSL/AQH wobbles and yet it's still quite obvious that something is unstable about even those numbers.
  •  Which are the outliers in the stats and which are reality? 
  •  Is a year enough time to allow to see?  
  •  How many years' averages will it take to know for certain?
Pity the poor media buyer and radio seller trying to make sense of them.

My conclusion, though I am not sure Arbitron (or BBM) would agree:  the Fresno sample isn't large enough, given the complexities of ethnic/non-ethnic sample balance, let alone Spanish only/bilingual/English only language and cell phone/land line home phone preference trends.

Nielsen just announced that it plans to double its TV diary and quadruple it's metered panel samples.

You'd think that would make a huge difference, and yet it may not since the TV ratings company in the USA did the same thing only five years ago with the goal of tripling it by 2011 and yet they seem to "need" to do it again now and state that they plan to move as fast to do so as their clients want.

Nielsen has competition which is clearly driving the changes; Arbitron does not. 

Same in Canada, where BBM is the only option for broadcasters, since they "own" it.

If North American radio wants to know how big samples must be to be consistent and credible in our markets + how fast our ratings supplier can get there, we're going to have to keep the pressure on ARB and BBM, with high hopes they care as much as their clients do, to make changes in spite that.

Thursday, February 16, 2012

36% Of 25-54 Country Radio Listeners Live In Cell Phone Only Homes

Top line findings from Albright & O'Malley's 7th Annual "Roadmap 2012" country format perceptual will be presented Tuesday, February 21, at our annual client seminar, but just one stat from the huge sample from 94 country stations (28% of them from all across Canada and 72% listeners to A&O's American country client stations) points to why in spite of their very positive financial results, reported yesterday, Arbitron (and Canada's BBM too) is having less positive results when it comes to delivering a consistent, quality sample.

The 25-54 composition of the 2012 A&O country format P-1 study is 27% 25-34, 31% 35-44 and 42% 45-54.

This year's country 25-54 CPO percentage is actually down slightly from last year's figure, which was 40%.

Thursday, February 02, 2012

ARB's Twin Waterbeds, Diary And PPM

RBR: "Given the news that the Media Rating Council (MRC) that it has withdrawn accreditation of the monthly AQH radio ratings data produced by the PPM service in Cleveland, Portland OR, Riverside-San Bernardino, Salt Lake City-Ogden-Provo, and Tampa-St. Petersburg-Clearwater, we revisited an interview with MRC Executive Director and CEO George Ivie as to issues surrounding accreditation in PPM markets. First of all, it’s a complex thing to accredit a new service and Arbitron has adjusted the service to meet the compliance needs of the MRC. That takes time and is always under review."

“It’s a complex problem. It’s like a waterbed, you push down on one side and other sides move. You have to push down on the whole bed to hold it all down. If you let go of s
omething, certain things may get worse, other things may get better. You have to really be vigilant.” -- Ivie in Inside Radio

Since the news broke, clients in both PPM and diary markets have been asking if "panic" is the order of the day, as Arbitron Director of Programming Services Jon Miller coincidentally titled his new blog post on ARB's training site and my advice given that, yes, credibility of media measurement is absolutely crucial, but since we've all seen this coming for a very long time is: DON'T (panic).

Keep the pressure on researchers to be as open as possible as they work on the complex sample quality issues that every market is witnessing. Secrecy won't improve anyone's confidence level.

A quick recap of past history as chronicled on this blog alone:
Let's get up to speed about it all at CRS. (click the banner below for a free invitation if you're not in a competitive situation with an A&O client)

<span class=ao-banner" height="150" width="180">

Monday, July 18, 2011

Track This Number, Especially If You're In A PPM Market

The average household size in diary sample markets across both the USA and Canada has trended almost exactly the same as the census average, roughly 2.6 persons per home for a very long time.

Over the five decades I have been studying diary samples, there have been individual weeks during BBM or ARB survey periods when the number of persons per household for the total sample and individual radio stations have wobbled wildly away from that average, but when tracking multiple weekly samples for eight to twelve weeks it almost always trends back up or down to the norm.

Not so in PPM markets, where one very large household who loves (or hates!) your radio station can stay in the sample for 24 months or longer. That small number of high occupant households has the potential to help or harm your audience estimates for several years, giving new importance to fully understanding reliability in statistical terms.

If your station hits new, never before attained in diaries, high numbers in PPM:

1. Celebrate. Enjoy. Savor it. You have a station that some people use a lot and that's great.
2. Track the average number of people per ethnic and non-ethnic household in the total sample and also in your station's panelists week by week.
3. Troll for more homes just like them in your marketing efforts.
4. Educate your owner and manager. Hope that they understand that Newton's laws also apply to stats as well as physical objects. What goes up goes down as well.

Sunday, December 06, 2009

The 2010 U.S. Census Count Will Drive Your Ratings For The Next Decade

Writer Lornet Turnbull really got me to thinking of THE question to talk about at this week's Arbitron Fly-In, arguably even more fundamental than sample sizes, diary measurement or PPM:

Where will you be on Census Day — living in your RV, couch surfing at your friends', squatting in your parents' basement? The U.S. Census Bureau is preparing to count the more than 308 million men, women and children living in the country April 1, 2010. With just 10 questions on next year's form, this would seem simple enough. Yet the count is likely to be not just the most costly but possibly one of the most difficult ever staged.

"We are studying a population that is harder to count than the 2000 population," -census director Robert Groves

Both ARB and Nielsen need reliable census data if we are to hope that their audience estimates in the next ten years will be accurate and representative.

What are we all going to do if any group in the general population feels that the 2010 Census is suspect?

Future budgets .. from advertising to even federal revenue sharing, depend on getting it right!

Tuesday, September 29, 2009

Mark Ramsey at NAB Radio Show: Radio Should Drop Arbitron

He's always provocative, and last Friday in Philadelphia he was even more so than usual. And, as a continuation of my frustrated post below, this video (click to watch it) makes some powerful assertions, worth turning my blog over to today.

Sunday, September 27, 2009

Reconciling Two New Arbitron Reports (...And, Worrying!)

Last week, ARB released the first Radio Today update in two years, and added lots of great information on who uses as many as 57 different radio formats, where, when and how.

Then, just a day later, the company in association with Coleman Research put out another report looking at America's Most Successful stations and formats in PPM.

Unfortunately, it appears to purport that all but perhaps ten of those formats aren't going to drive enough cume to do well in metered measurement. Their lower cumes aren't large enough to consistently penetrate the smaller PPM sample. The great TSL and high loyalty which drove their success in diary measurement no longer registers as important to doing well in PPM methodology.

Shares no longer count very much, we are now told, since in PPM the majorty of lower cume stations all cluster together at the bottom of the new ranker, separated by small fractions ... which comes as quite a shock to many sellers and buyers who have become accustomed for four decades to using that metric to value the medium.

I am now staring nostalgically at these format performance indications from diaries, fearful that future "Radio Today" reports - as more markets drop diaries and move to PPM - will exhibit the disappearance of many mid-rankers, as they go from viability to financial starvation (unless they can find a way to grow a lot more "cume" - read that "get more meters of their heavy users in the sample" - and fast!).

By and large, PPM is treating Country well, but even our most successful stations in PPM have more cume and get many more listening occasions from their heaviest-users than the average does. Is this because their brands are stronger? Their execution is better?

Or, that they just were lucky enough to get meters into the hands of their listeners?


The rub: in highly-ethnic markets, the format gets very little cume from all but non-ethnic households, which can have "cume" (or, more correctly, number of meters) ramifications for both us and minority broadcasters.

How's this for an idea?

Since Arbitron is now saying that it plans to also measure "engaged" listening as well as "exposure," - which so much of PPM cume is - simply do one diary-based survey per year with a robust sample in every market so we can all rely on the accuracy of the audience shares of the engaged audience we have all come to trust and buy from for decades, and then, for those who want more accountability, track the behavior of individual station and format users, without trying to pass that much smaller sample off as an accurate measure of the size of average quarter hour audiences of stations with less reach and more loyal listeners.

Who says we have to choose between measuring exposure and engagement?

Can't you measure and compare both for us, please, Arbitron?

.. Just as you did so well in the last week.

Thursday, July 02, 2009

And, Then, They Came For Us

As World War II era Pastor Martin Niemöller's prescient quote implies, it's human nature for a majority group which feels unaffected by even extremely-egregious unfairness to ignore the concerns of their neighbors until it comes knocking.

Thus, since in most PPM-rated markets the country format has been doing pretty well, it has been tempting to ignore the complaints from minority radio owners and simply bask in the glow of our format's huge cumes and share growth.

At May's BCAB convention, for example, BBM CEO Jim McCleod noted in pre-currency tests of PPM in Vancouver it appeared country listening will go from a seven share of diary tuning to potentially more than a nine.

Now, one month before fast-growing and increasingly-ethnic San Diego goes currency with PPM, it has been concerning to watch as ARB appears in rumored pre-currency data to be having trouble locating the same percentage of core country radio listeners for the PPM panel as they had in the diary sample.

Unless something changes fast, country's audience share in San Diego could dip from the sevens to the fives.

Questions: are the population estimates which drive sample proportionality, now nine years away from the last census, accurate? Or, will they reflect a major change in the estimates when the 2010 census rolls into the ARB data in two years? What will that say about today's shares? Will panel management over the next few weeks pick up more country ultra core which seems to be missing in the pre-currency data we've seen thus far? What impact will that have on San Diego country radio owners Lincoln Financial Media and Clear Channel revenue projections as they budget for 2010 right now?

Suddenly, what seemed to be the worries of others have come to the country format's door in at least this one market where non-ethnic households have become a minority group.

When former ARB SVP/Ratings Services Jay Guyther blogs from his new post at ROI Media Solutions that this same issue has been on the back burner for 15 years at Arbitron and says "...leave the FCC, Attorneys General, Congress, MRC, etc. out of it. Increase the panel sizes to a sufficient size and most of the other problems, in my opinion, will go away..." it is long past time for action to replace homeostasis.

Tuesday, June 30, 2009

Of Skarzynski, Kabrich And PPM Samples (Again)

As Arbitron CEO Michael Skarzynski “welcomes any opportunity to discuss the importance of electronic measurement, the effectiveness of the PPM technology, the value of the data it produces, and our disciplined approach to the deployment of the service...” and heads to the House Of Representatives for a committee hearing (Oversight and Government Reform) to defend the company's device, sample sizes and fairness to minorities, consultant Randy Kabrich follows up his 2008 study and shows once again that sample size is directly related to PPM success for any but major mass appeal stations.
“One of the reasons that certain formats do not perform well in PPM (certain niche formats like smooth jazz to Hispanic formats) is heavily tied to the amount of sample available. It is all the more critical for Arbitron to have a balanced sample – correctly balanced by age, gender, ethnicity, geography, and socio-economic factors.”

Pity the poor station with a signal problem in even a very small part of their metro survey area or an audience that lives/works in only specific zip + four neighborhoods.

Will ARB work with owners/managers/programmers of these stations to drive their coverage area with a "test meter" to verify that the encoding is being picked up?

What should they do if they find a metered household in a small area within a zip code where an isolated signal issue is going to cost a subscribing station or format lots of revenue for the next 18 to 24 monthly books while they wait for that home to cycle out of the sample?

As Randy says (above) and your humble scribe posits as a result: the only way to assure that any station - ethnic or not - with that problem to get a fair shake over the long-term is for ARB's sample proportionality to be at least three times better than it has ever been, given a sample one third the size of the diary in-tab.

Sunday, May 10, 2009

Beasley Broadcast Group President/COO Bruce Beasley Says What Arbitron Hasn't

Last week, he blamed poor Arbitron PPM panel management on his quarterly analyst earnings conference call as one reason why the company’s Philadelphia revenues (off 31%) have been down more than the market average (which slipped 21%).

It’s not a secret that country’s core target is 35-54 and largely non-ethnic living more in suburbs than in the center of cities, so just as ARB was working to placate unhappy ethnic stations by improving the proportionality of African-Americans and Hispanics and calling its sample “solid” (September 2008) it now admits (according to Beasley) under-indexing of the 35-54 demo in those same months last fall of WXTU’s ratings decline.

Beasley claimed that 92.5 XTU had been consistently in the market’s top six or seven during the first year and a half of pre-currency and then the initial currency PPM monthlies, leading the company to believe that PPM was going to be kind to country radio, but, as minority broadcasters pushed ARB to improve the index of their listeners’ in the PPM panel, for some reason, WXTU fell to the low to mid-teens in the rankers.

Every country manager and programmer needs to demand, hope for, ask for, beg for, pay for excellent sample balance each month, at the very least within 10% +/-, in our targets too.

The Beasley exec last week told investors he feels that Arbitron is now correcting the problem:
“Recently we’ve seen some progress on 35-54 sampling and expect the revenue trend at WXTU to improve.”

He also said he “expects that [Arbitron] will not repeat the issue that impacted us in Philadelphia..” in Miami, where the company also owns WKIS, the country station in another highly ethnic market.

Are the same ARB people who provided those "solid" assurances last fall the ones promising Beasley that things are now under control in their markets?

Ask any researcher: sampling is both art and science. As sample sizes get smaller, “panel management” is the name of the art which Arbitron must consistently master equitably or everyone’s business is going to suffer.

Monday, April 13, 2009

Arbitron's Dr. Ed Responds

My recent post "The Right Way To Sample" generated this reply, which I'm delighted to receive and share with you:

Thanks for the opportunity for “equal time” (no, we won’t go there) to the comments in your April Fool’s Day blog about Arbitron’s sampling.

First of all, thanks for the kudos on the cell phone only sampling that’s now underway in 151 markets with the rest of the markets (except Puerto Rico) on the way this Fall. We’ve worked very hard to bring this positive enhancement to fruition after much testing. Oh, and by the way, your readers should know that the market in question that had double the intab for African-American was actually only 50% over for one phase and is presently 25% over for two phases. As we always say, it’s a twelve week survey…let’s see how the full twelve weeks finish. We do like to set the record straight.

Now, let’s go on to the issue about single versus multiple person per household sampling and Ted Bolton’s piece from 1995. The concept is known as “probability of selection”.

Let’s start with the factual errors because these often become “urban myths” (or perhaps with your blog, they could be “country myths”). Arbitron does not select individuals from a list. Arbitron uses a random digit dial (RDD) telephone sample of residential landline phone numbers. Landline phones, with a few exceptions, are tied to households. While not perfect, you can generally assume one landline phone number per household.

Ted spent much of his piece writing about lists and quotas. Those of you who still have budget for a perceptual are probably using lists of names and setting quotas for different demos (or more accurately, your research company is doing it). While quota samples tend to be the norm for custom research, that doesn’t make it right, just expedient. You can’t use quota samples to tabulate radio ratings that are projectable. We can’t call a house and say “We’re looking only for a Male 18-24 year old”. To do valid radio ratings, you must use a probability-based sample.

But on to the main point. In simple terms, Ted was wrong. Multiple person per household (MPPH) sampling has long been an accepted methodology in media surveys, government surveys, and many other kinds of surveys. And in terms of “probability of selection”, an MPPH frame equalizes the probability of selection across demos while a single person per household (SPPH) sample distorts the probability selection.

Here’s an example of what I mean:

In an SPPH frame, the larger the household, the smaller the chance that any one individual in the household will be selected.

Single person per household sampling is used in many surveys and one reason is that it makes sense for studies conducted by phone. Consider the logistics of talking to one person in a household for an extended period (perhaps 15 to 20 minutes for a perceptual or callout) and then asking to speak to another person. Those of us who design phone surveys do our best to dissuade anyone who would ask us to survey more than one person in a household by phone; it just doesn’t work. As someone who worked at Birch (which used single person per household sampling for those of you who remember Birch ratings), it would have been an operational nightmare to measure more than one person per household.

In Arbitron’s case, nearly everyone is eligible to be in the survey (one exception is most of this blog’s readers who are in the media business). We sample at the household level. Even our new address-based method of finding cell phone only households is designed for the household level. And each household has a relatively equal chance of being selected assuming they have a landline phone. With the addition of the cell phone only frame, that opportunity extends to nearly all households.

Let’s take the discussion one step further. One could weight for this difference in probability of selection. Some people in our business hate the idea of weighting and others would have us weight for almost every variable (are you right handed or left handed?). In statistical terms, weighting reduces bias and increases variance. In simpler terms, weighting makes up for the potential that a group that has different radio listening habits is not represented at their level of the population (bias), but increases bounce in the estimates (variance). It’s a tradeoff and our view is that the large amount of variance (bounce) that would be created by weighting for household size would be a major negative for your ratings.

We’re quite comfortable with measuring everyone above a certain age (6+ for PPM, 12+ for diary) in the household and the need to defend the sampling method has long passed. And thanks again, Jaye, for the opportunity to be part of your blog.

-- Ed Cohen, Vice President-Research Policy and Communication, Arbitron, Columbia, MD (410-312-8592 - Ed.Cohen@arbitron.com

OK, who else wants to add a few cents' worth? Add a comment below or drop me an email.

Wednesday, April 01, 2009

The Right Way To Sample

It's very good to see Arbitron begin cell-phone only household sampling in an initial 151 diary markets for the Spring 2009 radio survey which starts today.

However, in an Arbitrends market whose numbers came out just yesterday, where 7% of the population is African American, 15% of the in-tab in this month's proportionality report sent to stations by ARB is African American. Meanwhile, almost all of the market's non-ethnic stations which target 25-54 trended down.

How long do we need to complain about sampling before real dramatic change happens?

Read this, from longtime researcher Ted Bolton and take a guess when it was written:

"There is a very disturbing Arbitron bias that needs to be changed. The bias is a straight forward problem with Arbitron sampling methodologies that can be changed, and must be changed.

THE ARBITRON EQUAL OPPORTUNITIES VIOLATION

Without getting too technical, in order to understand the Arbitron problem, you first
need to understand just the basics of sampling theory. Sampling theory is based upon the concept of equal opportunity and the equal probability of being selected for the sample. If the universe consisted of 1,000 people, then all 1,000 people would have EXACTLY the same chance of getting selected.

The problem begins with how Arbitron samples. Arbitron selects an individual from a random list and then sends diaries to all of the individuals within the family unit. That means if the respondent is married with four people in the household over the age of 12, that respondent would receive six diaries!

The respondent who is single and lives alone (or even lives with somebody else on a non-married basis) receives a total of one diary.

In this scenario, Arbitron has completely violated the concept of equal opportunity.

Family members are eating up the age quotas much to the dismay of the single/live alone respondents.


What if you were doing callout research and instead of sampling individual listeners, you decided to sample each member of the household you contacted. Would this make any sense whatsoever?

The answer is no, yet we live by an Arbitron system where this is the norm.


THE RIGHT WAY TO SAMPLE


Arbitron sampling should occur on an individual level, not on a family level. The only reason Arbitron samples on a family level is to reduce the costs of completing the survey.

Take it from a researcher who knows filling quotas on a family level is cheaper than filling quotas on an individual level. We all want our listeners to have as much of a chance of being selected or an Arbitron survey as does anyone else."


Believe it or not, the above paragraphs are from "Radio Trends," a newsletter published by Bolton in 1995.

It could have been, should have been, written yesterday. Which is why I reprint it here, and now.

Wednesday, November 19, 2008

I Can't Top Bob Michaels

.. for his experience and expertise as an independent handicapper of the race just starting between Nielsen/Cumulus, Arbitron and Eastlan, but the most interesting development to me that I'll be anxious to see more about: data emerging as Nielsen recruits by address, not telephone.

If this works, it won't be too long before ARB and BBM in Canada do it too, I'll bet.

Inside Radio's Frank Saxe talked yesterday to Cumulus COO John Dickey, who claims that it will enlist a larger number of panelists than Arbitron has used in condensed markets. “That’s going to reduce the ridiculous bounces and very high margin of errors which led to a lack of confidence in buying our medium. That alone is going to go along way to introducing a high level of credibility in our medium.”

By recruiting diarykeepers by address, not the telephone, he estimates Nielsen will enlarge the potential sample by up to 40% to include cell phone-only households and people with unlisted numbers. They’ve also committed to oversampling hard-to-reach demos like 18-34.

As with all surveys, there are a lot of moving parts, but it seems to me that this one, especially, will be an important metric to track very carefully and hopefully.

Tuesday, March 04, 2008

Why An Index Of 80 Or 90 Isn't Good Enough

While it's good to hear that ARB's Pierre Bouvard is back from his vacation fully-refreshed and enthusiastic for the radio ratings giant's sample improvement goals for the coming year (according to Inside Radio): “2008 is the year of 25-34s.”

After suffering through last Summer’s sagging PPM in-tab rates, IR's Frank Saxe reports he got a phone call from Bouvard yesterday, touring that Arbitron is increasingly confident it has solved the problem of how to hit 18-34 targets. Its performance among adults has been steadily improving, taking its biggest steps forward in recent months due to a focus on the 18-24 demo. It’s an effort that’s included higher sampling rates, richer incentives and more interaction between the company and its panelists.

Arbitron president of sales and marketing Pierre Bouvard says “We’re not done, but we have made a lot of progress in four months.” An upgrade to Arbitron’s software allows its research team to begin honing-in on the next target. Bouvard says “We said we were going to go on a jihad on 18-24s and we did. Now it’s onward an upwards -— 2008 is going to be the year of the 25-34s.”

Since last month they’ve done that by applying many of the same strategies and techniques used on 18-24s. Arbitron’s four-month focus on 18-24s has led to a 13% improvement in that Philadelphia demo — and it’s up 113% over one year ago. Bouvard says “It’s nothing short of spectacular. I’m feeling very good about how the panels are shaping up.”

Here's the problem: as you look at day-by-day, hour by hour and even minute by minute data in PPM, the sample is often four times as big as the daypart, day and even diary page sample. That's because so many diarykeepers 'forget' to write down at least half of their actual listening. The laws of statistical replication state that in order to double the accuracy of a probability sample, you must quadruple the sample size, so in one sense the minute by minute cume a daypart in PPM could potentially be more accurate even than weekly daypart data in the diary, which may be why it looks so reasonable and often passes "the gut test" (does it look believable?). ARB's PPM advocate John Snyder has been making this claim lately. He's a very smart guy, who makes very convincing Power Point slide shows.

However, the total panel sample is actually smaller than the random diary sample (to save money, of course) and a panel sample isn't exactly random either. Just because something looks 'right' most of the time doesn't mean that scientifically, statistically it is so.

I am hoping that ARB decides to start guaranteeing 100% of their targets for both PPM and diary markets in this year of the 25-34. No less than 100% accuracy is acceptable at a time when radio needs to convince advertisers that we are worth 100% of the money they spend on us, not 80 ot 90%.