Skip to main content

Posts

Showing posts with the label Quality Measures

What is Data Quality, and How to Enhance it in Research

  We often talk about “data quality” or “data integrity” when we are discussing the collection or analysis of one type of data or another. Yet, the definition of these terms might be unclear, or they may vary across different contexts. In any event, the terms are somewhat abstract -- which can make it difficult, in practice, to improve. That is, we need to know what we are describing with those terms, before we can improve them. Over the last two years, we have been developing a course on   Total Data Quality , soon to be available on Coursera. We start from an error classification scheme adopted by survey methodology many years ago. Known as the “Total Survey Error” perspective, it focuses on the classification of errors into measurement and representation dimensions. One goal of our course is to expand this classification scheme from survey data to other types of data. The figure shows the classification scheme as we have modified it to include both survey data and organic f...

Total Data Quality Update

 We have been working hard on applying the Total Survey Error (TSE) concept to hybrid data sources. That is, data that includes both designed and gathered data. We use the term "designed" for data that are designed for analysis. Gathered data, on the other hand, are not designed for analysis.  We find ourselves more and more relying on multiple sources of data, and wanted to bring our quality perspective to those problems. It feels to me like our survey experience with quality assessment is highly relevant for either hybrid data situations and for gathered data. TSE gives us a way to think through the issues. We have been offering a series of webinars on the the topic for the last few summers. We are working toward a larger course. More on that topic soon...

Total Data Quality

In an earlier post, I suggested that survey methodologists are "data quality specialists." Our focus on " total survey error " (TSE) is, in many ways, the central defining concept of our field. This focus on data quality could be an important contribution that survey methodologists make to the emerging field of data science. But in order to make that contribution, we may need to test the fit of the TSE concept on evaluations of non-survey data. One of the sources of error in surveys that we examine in surveys is "nonresponse." Does this concept apply to other sources of data? Certainly other sources of data having missing data. But nonresponse is a specific mechanism where we sample a unit and then request data, but the unit fails to supply the data. How does this concept apply to other sources of data? I wouldn't say that Twitter data suffer from "nonresponse" due to the fact that not everyone has a Twitter account or even that not every...

Predictions of Nonresponse Bias

One issue that we have been discussing is indicators for the risk of nonresponse bias. There are some indicators that use observed information (i.e. largely sampling frame data) to determine whether respondents and nonrespondents are similar. The R-Indicator is an example of this type of indicator. It's not the only one. There are several sample balance indicators. There is an implicit model that the observed characteristics are related to the survey data and controlling for them will, therefore, also control the potential for nonresponse bias. Another indicator uses the observed data, including the observed survey data, and a model to fill in the missing survey data. The goal here is to predict whether nonresponse bias is likely to occur. Here, the model is explicit. An issue that impacts either of these approaches is that if you are able to predict the survey variables with the sampling frame data, then why bother addressing imbalances on them during data collection? One answ...

Response Rates and Responsive Design

A recent article by Brick and Tourangeau re-examines the data from a paper by Groves and Peytcheva (2008). The original analyses from Groves and Peytcheva were based upon 959 estimates with known variables measured on 59 surveys with varying response rates. They found very little correlation between the response rate and the bias on those 959 estimates. Brick and Tourangeau view the problem as a multi-level problem of 59 clusters (i.e. surveys) of the 959 estimates. They created for each survey a composite score based on all the bias estimates from each survey. Their results were somewhat sensitive to how the composite score was created. They do present several different ways of doing this -- simple mean, mean weighted by sample size, mean weighted by the number of estimates. Each of these study-level composite bias scores is more correlated with the response rate. They conclude: "This strongly suggests that nonresponse bias is partly a function of study-level characteristics; ...

Centralization vs Local Control in Face-to-Face Surveys

A key question that face-to-face surveys must answer is how to balance local control against the need for centralized direction. This is an interesting issue to me. I've worked on face-to-face surveys for a long time now, and I have had discussion about this issue with many people. "Local control" means that interviewers make the key decisions about which cases to call and when to call them. They have local knowledge that helps them to optimize these decisions. For example. if they see people at home, they know that is a good time to make an attempts. They learn people's work schedules, etc. This has been the traditional practice. This may be because before computers, there was no other option. The "centralized" approach says that the central office can summarize the data across many call attempts, cases, and interviewers and come up with  an optimal policy. This centralized control might serve some quality purpose, as in our efforts here to promote more...

Goodhart's Law

I enjoy listening to the data skeptic podcast. It's a data science view of statistics, machine learning, etc. They recently discussed Goodhart's Law on the podcast. Goodhart's was an economist. The law that bears his name says that "when a measure becomes a target, then it ceases to be a good measure." People try and find a way to "game" the situation. They maximize the indicator but produce poor quality on other dimensions as a consequence. The classic example is a rat reduction program implemented by a government. They want to motivate the population to destroy rats, so they offer a fee for each rat that is killed. Rather than turn in the rat's body, they just ask for the tail. As a result, some persons decide to breed rats and cut off their tails. The end result... more rats. I have some mixed feelings about this issue. There are many optimization procedures that require some single measure which can be either maximized or minimized. I think th...

Balancing Response through Reduced Response Rates

A case can be made that balanced response -- that is, achieving similar response rates across all the subgroups that can be defined using sampling frame and paradata -- will improve the quality of survey data. A paper that I was co-author on used simulation with real survey data to show that actions that improved the balance of response usually led to reduced bias in adjusted estimates. I believe the case is an empirical one. We need more studies to speak more generally about how and when this might be true. On the other hand, I worry that studies that seek balance by reducing response rates (for high-responding groups) might create some issues. I see two types of problems. First, low response rates are generally easier to achieve. It takes skills and effort to achieve high response rates. The ability to obtain high response rates, like any muscle, might be lost if it is not used. Second, if these studies justify the lower response rate by saying that estimates are not significantly ...

Balancing response... without simply retreating

I've seen several studies that examine whether "balancing response" with respect to a set of covariates available on the frame can lead to reductions in nonresponse bias. Most of the studies indicate that more balanced response is associated with less nonresponse bias. However, there is a strategy for balancing response that worries me a bit -- reducing the response rates of the groups that have the highest response rates and, thereby, reducing the overall response rate. Why does this worry me? Several reasons. First, when does this work? We have some studies that show reductions in bias. The studies that show increases in bias might be suppressed due to publication bias. So, how are we supposed to know when it works and when it doesn't? Second, it's easy to reduce response rates. It's harder to raise them. What's worse, once we reduce response rates, how do we ever get back the skills required for obtaining higher response rates. Maybe we are simp...

Attrition in Designs that use Frequent Measurement

I saw this paper recently that talked about how to measure and evaluate nonresponse to surveys that use short, frequently-administered instruments ("measurement-burst survey"). I've been working on a problem with data like these for a while. A complication was that the questionnaire changed based upon the intervals between measurements. For example, questions might begin, "Since you last completed this survey..." or "in the last two weeks..." depending upon the situation. Plus, panel members could choose to respond at different intervals, even though they were asked to respond at a specified interval. This made for a complex pattern of missing data. I ended up defining attrition in several ways.  The most useful was to lay out a grid over time. The survey was designed to be taken weekly, so I looked at each week over the time period to see if any reporting occured. This allowed me to how many cells in the grid were missing. But even that wasn...

Context and Daily Surveys

I've been reading a very interesting book on daily diary surveys. One of the chapters, by Norbert Schwarz, makes some interesting points about how frequent measurement might not be the same as a one-time measurement of similar phenomena. Schwarz points to the well-known studies that he did where they varied the scale of measurement. One of the questions was about how much TV people watch. One scale had a maximum of something like 10 or more hours per week, while the other had a maximum of 2.5 hours per week. The reported distributions changed across the two different scales. It seems that people were taking normative cues from the scale, i.e. if 2.5 hours is a lot, "I must view less than that," or "I don't want to report that I watch that much TV when most other people are watching less." He points out that daily surveys may provide similar context clues about normative behavior. If you ask someone about depressive episodes every day, they may infer that...

Sensitivity Analysis and Nonresponse Bias

For a while now, when I talk about the risk of nonresponse bias, I suggest that researchers look at the problem from as many different angles as possible, employing varied assumptions. I've also pointed to work by Andridge and Little that uses proxy pattern-mixture models and a range of assumptions to do sensitivity analysis. In practice, these approaches have been rare. A couple of years ago, I saw a presentation at JSM that discussed a method for doing sensitivity analyses for binary outcomes in clinical trials with two treatments. The method they proposed was graphical and seemed like it would be simple to implement. An article on the topic has now come out. I like the idea and think it might have applications in surveys. All we need are binary outcomes where we are comparing two groups. It seems that there are plenty of those situations.

Web Panels vs Mall Intercepts

I saw this interesting article that just came out. It called to my mind a talk that was hosted here a few (8?) years ago. The talk was someone from a major corporation who talked about how they switched product testing from church basements to online panels. They found that once they switched, the data became worse. The online panels picked products that ended up failing at higher rates. This seemed like a tough problem. There isn't much of a "nonresponse" kind of relationship here. But at least understanding the mechanism that got people into online panels and how they were then selected and agreed to participate in this kind of product testing seemed important. It's not my area, so I'm wondering if this has ever been done. Not that anyone would understand the process of recruiting people to participate in product testing in church basements. But that process at least worked. This new article looks at an old process -- mall intercepts -- for recruiting people...

Probability Sampling

In light of the recent kerfuffle over probability versus non-probability sampling, I've been thinking about some of the issues involved with this distinction. Here are some thoughts that I use to order the discussion in my own head: 1. The research method has to be matched to the research question. This includes cost versus quality considerations. Focus groups are useful methods that are not typically recruited using probability methods. Non-probability sampling can provide useful data. Sometimes non-probability samples are called for. 2. A role for methodologists in the process is to test and improve faulty methods. Methodologists have been looking at errors due to nonresponse for a while. We have a lot of research for using models to reduce nonresponse bias. As research moves into new arenas, methodologists have a role to play there. While we may (er... sort of) understand how to adjust for nonresponse, do we know how to adjust for an unknown probability of getting into an on...

The Dual Criteria for a Useful Survey Design Feature

I've been working on a review of patterns of nonresponse to a large survey on which I worked. In my original plan, I looked at things that are related response, and then I looked at things that are related to key statistics produced by the survey. "Things" include design features (e.g. number of calls, refusal conversions, etc.) and paradata or sampling frame data (e.g. Census Region, interviewer observations about the sampled unit, etc.). We found that there were some things that heavily influenced response (e.g. calls) that did not influence the key statistics. Good, since more or less of that feature, although important for sampling error, doesn't seem important with respect to nonresponse bias. There were also some that influenced the key statistics but not response. For example, interviewer observations we have for the study. The response rates are close across subgroups of these estimates. As a result, I won't have to rely on large weights to get to unbi...

Better to Adjust with Weights, or Adjust Data Collection?

My feeling is that this is a big question facing our field. In my view, we need both of these to be successful. The argument runs something like this. If you are going to use those variables (frame data and paradata) for your nonresponse adjustments, then why bother using them to alter your data collection? Wouldn't it be cheaper to just use them in your adjustment strategy? There are several arguments that can be used when facing these kinds of questions. The main point I want to make here is that I believe that this is an empirical question. Let's call X my frame variable and Y the survey outcome variable. If I assume that the relationship between X and Y is the same no matter what the response rate for categories of X, then, sure, it might be cheaper to adjust. But that doesn't seem to be true very often. And that is an empirical question. There are two ways to examine this question. [Well, whenever someone says definitively there are "two ways of doing someth...

Covariates of Measurement Error

I've been working on some mixed-mode problems where nonresponse and measurement error are confounded. I recently read an interesting article on using adjustment models to disentangle the two sources of error. The article is by Vannieuwenhuyze, Loosveldt, and Molenberghs. They suggest that you can make adjustments for measurement error if you have things that predict when those errors occur. They give specific examples. It's things that measure social conformity and other hypothesized mechanisms that lead to response error. This was very interesting to read about. I suppose that just as with nonresponse, the predictors of this error -- in order to be useful -- need to predict when those errors occur and the survey outcome variables themselves. This is a new and difficult task... but one worth solving giving the push to use mixed mode designs.

Proxy Y's

My last post was a bit of crankiness about the term "nonresponse bias." There is a bit of terminology, on the other hand, that I do like -- "Proxy Y's." We used this term in a paper a while ago. The thing that I like about this term, is that it puts the focus on the prediction of Y. Based on the paper by Little and Vartivarian (2005), this seemed like a more useful thing to have. And we spent time looking for things that could fit the bill. If we have something like this, the difference between responders and the full sample might be a good proxy for bias with the actual Y's. I'm not backtracking here -- it's still not "nonresponse bias" in my book. It's just a proxy for it. The paper we wrote found that good proxy Y's are hard to find. Still, it's worth looking. And, as I said, the term keeps us focused on finding these elusive measures. 

When should we use the term "nonresponse bias"?

Maybe I'm just being cranky, but I'm starting to think we need to be more careful about when we use the term "nonresponse bias." It's a simple term, right? What could be wrong here? The situation that I'm thinking about is when we are comparing responders and nonresponders on characteristics that are known for everyone. This is a common technique. It's a good idea. Everyone should do this to evaluate the quality of the data. My issue is when we start to describe the differences between responders and nonresponders on these characteristics as "nonresponse bias." These differences are really proxies for nonresponse bias. We know the value for every case, so there isn't any nonresponse bias. The danger, as I see it, is that naive readers could miss that distinction. And I think it is an important distinction. If I say "I have found a method that reduces nonresponse bias," what will some folks hear? I think such a statement is pro...

Monitoring Daily Response Propensities

I've been working on this paper for a while. It compares models estimated in the middle of data collection with those estimated at the end of data collection. It points out that these daily models may be vulnerable to biased estimates akin to the "early vs. late" dichotomy that is sometimes used to evaluate the risk of nonresponse bias.The solution is finding the right prior specification in a Bayesian setup or using the right kind and amount of data from a prior survey so that estimates will have sufficient "late" responders. But, I did manage to manufacture this figure which shows the estimates from the model fit each day with the data available that day ("Daily") and the model fit at the end of data collection ("Final"). The daily model is overly optimistic early. For this survey, there were 1,477 interviews. The daily model predicted there would be 1,683. The final model predicted 1,477. That's the average "optimism." ...