Skip to main content

Posts

Showing posts with the label Response rates

Response Rates and Responsive Design

A recent article by Brick and Tourangeau re-examines the data from a paper by Groves and Peytcheva (2008). The original analyses from Groves and Peytcheva were based upon 959 estimates with known variables measured on 59 surveys with varying response rates. They found very little correlation between the response rate and the bias on those 959 estimates. Brick and Tourangeau view the problem as a multi-level problem of 59 clusters (i.e. surveys) of the 959 estimates. They created for each survey a composite score based on all the bias estimates from each survey. Their results were somewhat sensitive to how the composite score was created. They do present several different ways of doing this -- simple mean, mean weighted by sample size, mean weighted by the number of estimates. Each of these study-level composite bias scores is more correlated with the response rate. They conclude: "This strongly suggests that nonresponse bias is partly a function of study-level characteristics; ...

Every Hard-to-Interview Respondent is Difficult in their Own Way...

The title of this post is a paraphrase of a saying coined by Tolstoi. " Happy families are all alike; every unhappy family is unhappy in its own way." I'm stealing the concept to think about survey respondents.  To simplify discussion, I'll focus on two extremes. Some people are easy respondents. No matter what we do, no matter how poorly conceived, they will respond. Other people are difficult respondents. I would argue that these latter respondents are heterogenous with respect to the impact of different survey designs on them. That is, they might be more likely to respond under one design relative to another. Further, the most effective design will vary from person to person within this difficult group.  It sounds simple enough, but we don't often carry this idea into practice. For example, we often estimate a single response propensity, label a subset with low estimated propensities as difficult, and then give them all some extra thing (often more money). ...

Slowly Declining Response Rates are the Worst!

I have seen this issue on several different projects. So I'm not calling out anyone in particular. I keep running into this issue. Repeated cross-sectional surveys are the most glaring example, but I think it happens other places as well. The issue is that with a slow decline, it's difficult to diagnose the source of the problem. If everything is just a little bit more difficult (i.e. if contacting persons, convincing people to list a household, finding the selected person, convincing them to do the survey, and so on), then it's difficult to identify solutions. One issue that this sometimes creates is that we keep adding a little more effort each time to try to counteract the decline. A few additional more calls. A slightly longer field period. We don't then search for qualitatively different solutions. That's not to say that we shouldn't make the small changes. Rather, that they might need to be combined with longer term planning for larger changes. That...

Reasons for maintaining high response rates

A few years ago, I was presenting at a conference of substantive experts. I gave an update on a progress on a survey of interest to this group. I talked about how nonresponse bias can be complex, and that the response rate might not be a good predictor of when this bias occurs -- based on Groves and Peytcheva . I was speaking with one of the researchers after my presentation, and I was surprised to hear her say that she interpreted my comments to mean that "response rates don't matter." Although that interpretation makes sense, it hadn't really occurred to me in that way until she said it. Since then, it seems like we've seen a lot of published papers and conference presentation where lowering the response rate becomes a tactic for improving the survey. Most studies taking this tactic lower the response rates for groups that tend to respond at higher rates. The purported benefit is  response set balance on known characteristics from the sampling frame is improve...

Balancing Response through Reduced Response Rates

A case can be made that balanced response -- that is, achieving similar response rates across all the subgroups that can be defined using sampling frame and paradata -- will improve the quality of survey data. A paper that I was co-author on used simulation with real survey data to show that actions that improved the balance of response usually led to reduced bias in adjusted estimates. I believe the case is an empirical one. We need more studies to speak more generally about how and when this might be true. On the other hand, I worry that studies that seek balance by reducing response rates (for high-responding groups) might create some issues. I see two types of problems. First, low response rates are generally easier to achieve. It takes skills and effort to achieve high response rates. The ability to obtain high response rates, like any muscle, might be lost if it is not used. Second, if these studies justify the lower response rate by saying that estimates are not significantly ...

Mode Sequence

A few years ago, I did an experiment with two sequences of modes for a screening survey. The modes were mail and face-to-face. We found that the sequence didn't matter much for the response rate to the screener, but that the arm that started with face-to-face and then used mail had a better response rate to the main interview given to those who were found to be eligible in the screening interview. There are other experiments that use different sequences of modes. Some of these find that the sequence doesn't matter. For example, Dillman and colleagues looked at mail-telephone and telephone-mail and these had about the same response rate. On the other hand, Millar and Dillman found that for mail-web mixed-mode surveys the sequence does seem to matter, although certainly the number and kind of contact attempts are also important. It does seem that there are times when the early attempts might interfere with the effectiveness of later attempts. That is, we "harden the ref...

Is the "long survey" dead?

A colleague sent me a link to a blog arguing that the "long survey" is dead. The blog takes the point of view that anything over 20 minutes is long. There's also a link to another blog that presents data from survey monkey surveys showing that the longer the questionnaire, the less time that is spent on each question. They don't really control for question length, etc. But it's still suggestive. In my world 20 minutes is still a short survey. But the point is still taken. There has been some research on the effect of survey length (announced) on response rates. There probably is need for more. Still, it might be time to start thinking of alternatives to improve response to long surveys. The most common is to offer a higher incentive, and thereby counteract the burden of the longer survey. Another alternative is to shorten the survey. This doesn't work if your questions are the ones getting tossed. Of course, substituting big data for elements of surveys is...

Responsive Design and Quota Sampling

I conducted a webinar on responsive design this week. I had several interesting questions. One of these was a question about responsive design and quota sampling.  The question was whether these two approaches are, in fact, different? Of course, there are similarities in that the response process is being controlled -- somewhat -- by the researchers. And this may lead to "allocating" nonresponse to some groups over others. For example, if some group is responding at higher rates, we might allocate resources to the lower responding group. Quota sampling will stop data collection for groups that have reached their quota. There are differences, however. Responsive design attempts to provide balanced response, but doesn't necessarily force that to happen. Further, responsive design is attempting to control the data collection process using a variety of approaches. Quota sampling only has one approach -- stop when the quota is full.  I do worry that there may be a conver...

Margin of Error

There was a debate held yesterday on "Margin of Error" in the presence of nonresponse and using non-probability samples. This is an interesting and useful discussion. In the best of circumstances, "margin of error" represents the sampling error associated with an estimate. Unfortunately, other matters often... errrrr... always interfere. The sampling mechansim is not easily identified or modeled in the case of nonprobability samples. In the case of probability samples, the nonresponse mechanism has to be modeled. Either of these situations involve some model assumptions (untestable) that are required to motivate the estimation of a margin of error. One step forward would be for people who report estimated "margins of error" to reveal all of their assumptions in their  models (weighting models for nonresponse or, in the case of nonprobability samples, selection) and describe the sampling and recruitment mechanisms sufficiently such that others can evalu...

Sensitivity Analysis and Nonresponse Bias

For a while now, when I talk about the risk of nonresponse bias, I suggest that researchers look at the problem from as many different angles as possible, employing varied assumptions. I've also pointed to work by Andridge and Little that uses proxy pattern-mixture models and a range of assumptions to do sensitivity analysis. In practice, these approaches have been rare. A couple of years ago, I saw a presentation at JSM that discussed a method for doing sensitivity analyses for binary outcomes in clinical trials with two treatments. The method they proposed was graphical and seemed like it would be simple to implement. An article on the topic has now come out. I like the idea and think it might have applications in surveys. All we need are binary outcomes where we are comparing two groups. It seems that there are plenty of those situations.

Probability Sampling

In light of the recent kerfuffle over probability versus non-probability sampling, I've been thinking about some of the issues involved with this distinction. Here are some thoughts that I use to order the discussion in my own head: 1. The research method has to be matched to the research question. This includes cost versus quality considerations. Focus groups are useful methods that are not typically recruited using probability methods. Non-probability sampling can provide useful data. Sometimes non-probability samples are called for. 2. A role for methodologists in the process is to test and improve faulty methods. Methodologists have been looking at errors due to nonresponse for a while. We have a lot of research for using models to reduce nonresponse bias. As research moves into new arenas, methodologists have a role to play there. While we may (er... sort of) understand how to adjust for nonresponse, do we know how to adjust for an unknown probability of getting into an on...

The Dual Criteria for a Useful Survey Design Feature

I've been working on a review of patterns of nonresponse to a large survey on which I worked. In my original plan, I looked at things that are related response, and then I looked at things that are related to key statistics produced by the survey. "Things" include design features (e.g. number of calls, refusal conversions, etc.) and paradata or sampling frame data (e.g. Census Region, interviewer observations about the sampled unit, etc.). We found that there were some things that heavily influenced response (e.g. calls) that did not influence the key statistics. Good, since more or less of that feature, although important for sampling error, doesn't seem important with respect to nonresponse bias. There were also some that influenced the key statistics but not response. For example, interviewer observations we have for the study. The response rates are close across subgroups of these estimates. As a result, I won't have to rely on large weights to get to unbi...

Proxy Y's

My last post was a bit of crankiness about the term "nonresponse bias." There is a bit of terminology, on the other hand, that I do like -- "Proxy Y's." We used this term in a paper a while ago. The thing that I like about this term, is that it puts the focus on the prediction of Y. Based on the paper by Little and Vartivarian (2005), this seemed like a more useful thing to have. And we spent time looking for things that could fit the bill. If we have something like this, the difference between responders and the full sample might be a good proxy for bias with the actual Y's. I'm not backtracking here -- it's still not "nonresponse bias" in my book. It's just a proxy for it. The paper we wrote found that good proxy Y's are hard to find. Still, it's worth looking. And, as I said, the term keeps us focused on finding these elusive measures. 

When should we use the term "nonresponse bias"?

Maybe I'm just being cranky, but I'm starting to think we need to be more careful about when we use the term "nonresponse bias." It's a simple term, right? What could be wrong here? The situation that I'm thinking about is when we are comparing responders and nonresponders on characteristics that are known for everyone. This is a common technique. It's a good idea. Everyone should do this to evaluate the quality of the data. My issue is when we start to describe the differences between responders and nonresponders on these characteristics as "nonresponse bias." These differences are really proxies for nonresponse bias. We know the value for every case, so there isn't any nonresponse bias. The danger, as I see it, is that naive readers could miss that distinction. And I think it is an important distinction. If I say "I have found a method that reduces nonresponse bias," what will some folks hear? I think such a statement is pro...

More methods research for the sake of methods...

In my last post, I suggested that it might be nice to try multiple survey requests on the same person. It reminded me of a paper I read a few years back on response propensity models that suggested continuing calling after the interview is complete, just so that you can estimate the model. At the time, I thought it was sort of humorous to suggest that. Now I'm drawing closer to that position. Not for every survey, but it would be interesting to try. In addition to validating estimated propensities at the person level, this might be another way to assess predictors of nonresponse that we can't normally assess. Peter Lugtig has an interesting paper and blog post about assessing the impact of personality traits on panel attrition. He suggests that nonresponse to a one-time, cross-sectional survey might have a different relationship to personality traits. Such a model could be estimated for a cross-sectional survey of employees who all have taken a personality test. You could do...

Estimating Response Probabilities for Surveys

I recently went to a workshop on adaptive treatment regimes. We were presented with a situation where they were attempting to learn about the effectiveness of a treatment to help with a chronic condition like addiction to smoking. The treatment is applied at several points over time, and can be changed based on changes in the condition of the person (e.g. they report stronger urges to smoke). In this setup, they can learn effective treatments at the patient level. In surveys, we only observe successful outcomes one time. We get the interview, we are done. We estimate response propensities by averaging over sets of cases. Within in any set, we assume that each person is exchangeable. Not by observing response to multiple survey requests on the same person. Even panel surveys are only a little different. The follow-up interviews are often only with cases that responded at t=1. Even when there is follow-up with the entire sample, we usually leverage the fact that this is follow-up to ...

Are we really trying to maximize response rates?

I sometimes speculate that we may be in a situation where the following is true: Our goal is to maximize response rate We research methods to do this We design surveys based on this Of course, the real world is never so "pure." I'm sure there must be departures from this all the time. Still, I wonder what the consequences of maximizing (or minimizing) something else would be. Could research on increasing response still be useful under a new guiding indicator? I think that in order for older research to be useful under a new guiding indicator, the information about response has to be linked to some kind of subgroups in the sample. Indicators other than the response rate would place different values on each case (the response rate places the same value on each case). So for methods to be useful in a new world governed by some other indicator, those methods would have to useful for targeting some cases. On the simplest level, we don't want the average effect on ...

Speaking of costs...

I found another interesting article that talked about costs. This one , from Teitler and colleagues, described the apparent nonresponse biases present at different levels of cost per interview. This cuts to the chase on the problem. The basic conclusion was that, at least in this case, the most expensive interviews didn't change estimates. This enables discussing the tradeoffs in a more specific way. With a known amount of the budget that didn't prove to change estimates, could you make greater improvements by getting more cases that cost less? Spending more on questionnaire design? etc. Of course, that's easy to say after the fact. Before the fact, armed with less than complete knowledge, one might want to go after the expensive cases to be sure they are not different. Of course, I'd argue that you'd want to do that in a way that controlled costs (subsampling) until you achieve more certainty about the value of those data.

New Objective Functions...

I've argued in previous posts that the response rate has functioned like an objective function that has been used to design "optimal" data collections. The process has been implicitly defined this way. And it is probably the case that the designs are less than optimal for maximizing the response rate. Still, data collection strategies have been shaped by this objective function. Switching to new functions may be difficult for a number of reasons. First, we need other objective functions. These are difficult to define as there is always uncertainty with respect to nonresponse bias. Which function may be the most useful? R-Indicators? Functions of the relationships between observed Y's and sampling frame data? There are theoretical considerations, but we also need empirical tests. What happens empirically when data collection has a different goal? We haven't systematically tested these other options and their impact on the quality of the data. That should be hig...

Adjusting with Weights... or Adjusting the Data Collection?

I just got back from JSM where I saw some presentations on responsive/adaptive design. The discussant did a great job summarizing the issues. He raised one of the key questions that always seems to come up for these kinds of designs: If you have those data available for all the cases, why bother changing the data collection when you can just use nonresponse adjustments to account for differences along those dimensions? This is a big question for these methods. I think there are at least two responses (let me know if you have others). First, in order for those nonresponse adjustments to be effective, and assuming that we will use weighting cell adjustments (the idea extends easily to propensity modeling), the respondents within any cell need to be equivalent to a random sample of the cell. That is, the respondents and nonrespondents need to have the same mean for the survey variable. A question might be, at what point does that assumption become true? Of course, we don't know. B...