Skip to main content

Posts

Showing posts with the label Paradata

Should exceptions be allowed in survey protocol implementation?

I used to work on a CATI system (DOS-based) that allowed supervisors to release cases for calling through an override mechanism. That is, the calling algorithm had certain rules that kept cases out of the calling queue at certain times. The main thing was if something had been called and was a "ring-no-answer," then the system wouldn't allow it to be called (i.e. placed in the calling queue) until 4 hours had passed. But supervisors could override this and release cases for calling on a case-by-case basis. This was handy -- when sample ran out, supervisors could release more cases that didn't fall within the calling parameters. This kept interviewers busy dialing. Recently, I've started to think about the other side of such practices. That is, it is more difficult to specify the protocol that should be applied when these exceptions are allowed. Obviously, if the protocol is not calling a case less than four hours after a ring-no-answer, then the software explicit...

Responsive design and sampling variability

At the Joint Statistical Meetings, I went to a session on responsive and adaptive design. One of the speakers, Barry Schouten, contrasted responsive and adaptive designs. One of the contrasts was that responsive design was concerned with controlling short-term fluctuations in outcomes such as response rates. This got me thinking. I think the idea is that responsive design will respond to the current data, which includes some sampling error. In fact, it's possible that sampling error could be the sole driver of responsive design interventions in some cases. I don't think this is usually the case, but it certainly is part of what responsive designs might do. At first, this seemed like a bad feature. One could imagine that all responsive design interventions should include a feature that accounts for sampling error. For instance, decision rules that attain a level of statistical significance. We've implemented some like that. On the other hand, sometimes controlling samp...

Every Hard-to-Interview Respondent is Difficult in their Own Way...

The title of this post is a paraphrase of a saying coined by Tolstoi. " Happy families are all alike; every unhappy family is unhappy in its own way." I'm stealing the concept to think about survey respondents.  To simplify discussion, I'll focus on two extremes. Some people are easy respondents. No matter what we do, no matter how poorly conceived, they will respond. Other people are difficult respondents. I would argue that these latter respondents are heterogenous with respect to the impact of different survey designs on them. That is, they might be more likely to respond under one design relative to another. Further, the most effective design will vary from person to person within this difficult group.  It sounds simple enough, but we don't often carry this idea into practice. For example, we often estimate a single response propensity, label a subset with low estimated propensities as difficult, and then give them all some extra thing (often more money). ...

The Cost of a Call Attempt

We recently did an experiment with incentives on a face-to-face survey. As one aspect of the evaluation of the experiment, we looked at the costs associated with each treatment (i.e. different incentive amounts). The costs are a bit complicated to parse out. The incentive amount is easy, but the interviewer time is hard. Interviewers record their time for at the day level, not at the housing unit level. So it's difficult to determine how much a call attempt costs. Even if we had accurate data on the time spent making the call attempt, there would still be all the travel time from the interviewer's home to the area segment. If I could accurately calculate that, how would I spread it across the cost of call attempts? This might not matter if all I'm interested in is calculating the marginal cost of adding an attempt to a visit to an area segment. But if I want to evaluate a treatment -- like the incentive experiment -- I need to account for all the interviewer costs, as b...

Messy Experiments

I have this feeling that survey experiments are often very messy. Maybe it's just in comparison to the ideal type -- a laboratory with a completely controlled environment where only one variable is altered between two randomly assigned groups. But still, surveys have a very complicated structure. We often call this the "essential survey conditions." But that glib phrase might hide some important details. My concern is that when we focus on a single feature of a survey design, e.g. incentives, we might come to the wrong conclusion if we don't consider how that feature interacts with other design features. This matters when we attempt to generalize from published research to another situation. If we only focus on a single feature, we might come to the wrong conclusion. Take the well-known result -- incentives work! Except that the impact of incentives seems to be different for interviewer-administered surveys than for self-administered surveys. The other features of...

Goodhart's Law

I enjoy listening to the data skeptic podcast. It's a data science view of statistics, machine learning, etc. They recently discussed Goodhart's Law on the podcast. Goodhart's was an economist. The law that bears his name says that "when a measure becomes a target, then it ceases to be a good measure." People try and find a way to "game" the situation. They maximize the indicator but produce poor quality on other dimensions as a consequence. The classic example is a rat reduction program implemented by a government. They want to motivate the population to destroy rats, so they offer a fee for each rat that is killed. Rather than turn in the rat's body, they just ask for the tail. As a result, some persons decide to breed rats and cut off their tails. The end result... more rats. I have some mixed feelings about this issue. There are many optimization procedures that require some single measure which can be either maximized or minimized. I think th...

Balancing Response through Reduced Response Rates

A case can be made that balanced response -- that is, achieving similar response rates across all the subgroups that can be defined using sampling frame and paradata -- will improve the quality of survey data. A paper that I was co-author on used simulation with real survey data to show that actions that improved the balance of response usually led to reduced bias in adjusted estimates. I believe the case is an empirical one. We need more studies to speak more generally about how and when this might be true. On the other hand, I worry that studies that seek balance by reducing response rates (for high-responding groups) might create some issues. I see two types of problems. First, low response rates are generally easier to achieve. It takes skills and effort to achieve high response rates. The ability to obtain high response rates, like any muscle, might be lost if it is not used. Second, if these studies justify the lower response rate by saying that estimates are not significantly ...

What is a "response propensity"?

We talk a lot about response propensities. I'm starting to think we actually create a lot of confusion for ourselves by the way we sometimes have these discussions. First, there is a distinction between an actual and an estimated propensity. This distinction is important as our models are almost always misspecified. It is probably the case that important predictors are never observed -- for example, the mental state of the sampled person at the moment that we happen to contact them. So that the estimated propensity and true propensity are different things. The model selection choices we make can, therefore, have something of an arbitrary flavor to them. I think the choices we make should depend on the purpose of the model. We examined in a recent paper on nonresponse weighting whether call record information, especially the number of calls and refusal indicators, were useful predictors of response propensities for this purpose. It turns out that these variables were strong predic...

Survey Data and Big Data

I had an opportunity to revisit an article by Burns and colleagues that looks at using data from smartphones (they have a nice appendix of all the data they can get from each phone) to predict things that might trigger episodes of depression. Of course, the data don't contain any specific measures of depression. In order to get those, the researchers had to.... surveys. Once they had those, then they could find the associations with the censor data from the phone. Then they could deliver interventions through the phone. There are 38 sensors on the phone. The phone delivers data quite frequently. So even a small number of phones (n=8 in this trial) there was quite a large amount of data generated. A bigger trial would have even more data. So this seems like a big data application. And, in this case the "organic" data from the phone need some "designed" (i.e. survey) data in order to be useful. This is also interesting in that the smartphone is delivering a...

Training for Paradata

Paradata are messy data. I've been working with paradata for a number of years, and find that there are all kinds of issues. The data aren't always designed with the analyst in mind. They are usually a by-product of a process. The interviewers aren't focused (and rightly so) on generating high-quality paradata. In many situations, they sacrifice the quality of the paradata in order to obtain an interview. The good thing about paradata is that analysis of paradata is usually done in order to inform specific decisions. How should we design the next survey? What is the problem with this survey? The analysis is effective if the decisions seem correct in retrospect. That is, if the predictions generated by the analysis lead to good decisions. If students were interested in learning about paradata analysis, then I would suggest that they gain exposure to methods in statistics, machine learning, operations research, and an emerging category "data science." It seems ...

WSS Mini-Conference on Paradata

Next week, after the big storm, the Washington Statistical Society is sponsoring a mini-conference: "Benefits and Challenges in Using Paradata." The program is available online. This will be a nice opportunity to meet and discuss with folks working on similar problems. We are few in number. It's good to take advantage of these opportunities. I'm going to be speaking about problems with working with incoming streams of paradata. I can propose some solutions, but we need to get better at this.

Bayesian Adaptive Survey Design

Just a short blog post. I recently attended the 4th Workshop on Adaptive and Responsive Survey Design . There were many good papers delivered at this workshop. There was a particular focus on Bayesian approaches to the estimation of survey design parameters or paradata modeling. The link has some of the slides and papers.

Balancing response... without simply retreating

I've seen several studies that examine whether "balancing response" with respect to a set of covariates available on the frame can lead to reductions in nonresponse bias. Most of the studies indicate that more balanced response is associated with less nonresponse bias. However, there is a strategy for balancing response that worries me a bit -- reducing the response rates of the groups that have the highest response rates and, thereby, reducing the overall response rate. Why does this worry me? Several reasons. First, when does this work? We have some studies that show reductions in bias. The studies that show increases in bias might be suppressed due to publication bias. So, how are we supposed to know when it works and when it doesn't? Second, it's easy to reduce response rates. It's harder to raise them. What's worse, once we reduce response rates, how do we ever get back the skills required for obtaining higher response rates. Maybe we are simp...

Attrition in Designs that use Frequent Measurement

I saw this paper recently that talked about how to measure and evaluate nonresponse to surveys that use short, frequently-administered instruments ("measurement-burst survey"). I've been working on a problem with data like these for a while. A complication was that the questionnaire changed based upon the intervals between measurements. For example, questions might begin, "Since you last completed this survey..." or "in the last two weeks..." depending upon the situation. Plus, panel members could choose to respond at different intervals, even though they were asked to respond at a specified interval. This made for a complex pattern of missing data. I ended up defining attrition in several ways.  The most useful was to lay out a grid over time. The survey was designed to be taken weekly, so I looked at each week over the time period to see if any reporting occured. This allowed me to how many cells in the grid were missing. But even that wasn...

Responsive Design and Surveys with Short Time Frames

Another interesting question that I had during the webinar that I recently gave concerned responsive design and surveys with short time frames. I have to say, I mostly work on surveys with relatively long time frames. The shortest data collection that I have worked on in the last few years is about one month. That's not to say that I think responsive design is not relevant for surveys with short field periods. I think it is. If anything, following the prescribed regimen may be more important. A key aspect of responsive design, in my mind, is that the process is pre-planned. The indicators that are monitored, the decision rules for implementing interventions, the interventions, all have to be pre-planned. In a short survey, this is particularly important as their isn't time for developing ad hoc solutions. In a former life, I worked on surveys that had field periods of a day or two. In those studies, there wouldn't have been time to meet, discuss, and decide. Given the s...

Adaptive Designs and Incentives

I've been working on a paper about an incentive experiment that we did. It raised some interesting issues. And made me recall one of my favorite papers. Trussell and Lavrakas looked at incentives to a follow-up survey. They found that if someone had refused or been difficult to contact in the initial, screening survey, then a higher incentive was needed than for someone who had not refused or been difficult to contact. The incentives they recommend also differed by some demographic characteristics as well. I liked this example since the adaptation was linked, in part, to the paradata. These are the kinds of adaptations I have the most interest in. They require learning on the part of the survey organization that happens during data collection. I have the feeling that these kinds of adaptations can be particularly powerful since in models predicting response, it is often the case that paradata overwhelm the predictive power of demographic characteristics. There are all sorts of...

Interviewer Travel and New Forms of Data

The Director of the Census Bureau, John Thompson, recently blogged about a field test for the 2020 Decennial Census Nonresponse Follow-up. They are testing a number of new features, including the use of smartphones in data collection. I've been working with GPS data from smartphones used by field interviewers. The data are complex, but may offer new insights into interviewer travel. Think of travel as a broad concept -- it's not just an expense or efficiency issue. The order in which calls are made may also relate to field outcomes like contact and response rates. Perhaps these GPS data can help us understand how interviewers currently make decisions about how to work their sample. For example, do they move past sampled housing units when they first arrive to the area segment? Is this action associated with higher contact rates? Of course, travel is also an expense or efficiency issue. I wouldn't want pushing for more efficient travel to interfere with other aspects o...

Quantity becomes Quality

A big question facing our field is whether it is better to adjust data collection or do post-data collection adjustments to the data in order to reduce nonresponse bias. I blogged about this a few months ago. In my view, we need to do both. I'm not sure how the argument goes that says we only need to adjust at the end. I'd like to hear more of that. In my mind, it must be an assumption that once you condition on the frame data, the biases disappear and that assumption is valid at all points during the data collection. That must be a caricature -- which is why I'd like to hear more of the argument from a proponent of the view. In my mind, that assumption may or may not be true. That's an empirical question. But it seems likely that at some point in the process of collecting data, particularly early on, that assumption is not true. That is, the data are NMAR, even when I condition on all my covariates (sampling frame and paradata). Put another way, in a cell adjustmen...

The Dual Criteria for a Useful Survey Design Feature

I've been working on a review of patterns of nonresponse to a large survey on which I worked. In my original plan, I looked at things that are related response, and then I looked at things that are related to key statistics produced by the survey. "Things" include design features (e.g. number of calls, refusal conversions, etc.) and paradata or sampling frame data (e.g. Census Region, interviewer observations about the sampled unit, etc.). We found that there were some things that heavily influenced response (e.g. calls) that did not influence the key statistics. Good, since more or less of that feature, although important for sampling error, doesn't seem important with respect to nonresponse bias. There were also some that influenced the key statistics but not response. For example, interviewer observations we have for the study. The response rates are close across subgroups of these estimates. As a result, I won't have to rely on large weights to get to unbi...

"Go Big, or Go Home."

I just got back from JSM, where I participated in a session on adaptive design. Mick Couper served as a discussant for the session. The title of this blog post is one of the points from his talk. He said that innovative, adaptive methods need to show substantial results. Otherwise, it won't be convincing. As he pointed out, part of the problem is that we are often tinkering with marginal changes on existing surveys. These kinds of changes need to be low risk, that is, they can't cause damage to the results and should only help. However, these kinds of changes are often limited in what they can do. His point was to make some big changes that will show big effects may require some risk. This made sense to me. It would be nice to have some methodological studies that aren't constrained by the needs of an existing survey. I suppose this could be a separate, large sample with the same content as an existing survey. However, I wonder if this is a chicken or egg type of problem....