Skip to main content

Posts

Showing posts with the label Operational Issues

Centralization vs Local Control in Face-to-Face Surveys

A key question that face-to-face surveys must answer is how to balance local control against the need for centralized direction. This is an interesting issue to me. I've worked on face-to-face surveys for a long time now, and I have had discussion about this issue with many people. "Local control" means that interviewers make the key decisions about which cases to call and when to call them. They have local knowledge that helps them to optimize these decisions. For example. if they see people at home, they know that is a good time to make an attempts. They learn people's work schedules, etc. This has been the traditional practice. This may be because before computers, there was no other option. The "centralized" approach says that the central office can summarize the data across many call attempts, cases, and interviewers and come up with  an optimal policy. This centralized control might serve some quality purpose, as in our efforts here to promote more...

The Cost of a Call Attempt

We recently did an experiment with incentives on a face-to-face survey. As one aspect of the evaluation of the experiment, we looked at the costs associated with each treatment (i.e. different incentive amounts). The costs are a bit complicated to parse out. The incentive amount is easy, but the interviewer time is hard. Interviewers record their time for at the day level, not at the housing unit level. So it's difficult to determine how much a call attempt costs. Even if we had accurate data on the time spent making the call attempt, there would still be all the travel time from the interviewer's home to the area segment. If I could accurately calculate that, how would I spread it across the cost of call attempts? This might not matter if all I'm interested in is calculating the marginal cost of adding an attempt to a visit to an area segment. But if I want to evaluate a treatment -- like the incentive experiment -- I need to account for all the interviewer costs, as b...

Reasons for maintaining high response rates

A few years ago, I was presenting at a conference of substantive experts. I gave an update on a progress on a survey of interest to this group. I talked about how nonresponse bias can be complex, and that the response rate might not be a good predictor of when this bias occurs -- based on Groves and Peytcheva . I was speaking with one of the researchers after my presentation, and I was surprised to hear her say that she interpreted my comments to mean that "response rates don't matter." Although that interpretation makes sense, it hadn't really occurred to me in that way until she said it. Since then, it seems like we've seen a lot of published papers and conference presentation where lowering the response rate becomes a tactic for improving the survey. Most studies taking this tactic lower the response rates for groups that tend to respond at higher rates. The purported benefit is  response set balance on known characteristics from the sampling frame is improve...

Training for Paradata

Paradata are messy data. I've been working with paradata for a number of years, and find that there are all kinds of issues. The data aren't always designed with the analyst in mind. They are usually a by-product of a process. The interviewers aren't focused (and rightly so) on generating high-quality paradata. In many situations, they sacrifice the quality of the paradata in order to obtain an interview. The good thing about paradata is that analysis of paradata is usually done in order to inform specific decisions. How should we design the next survey? What is the problem with this survey? The analysis is effective if the decisions seem correct in retrospect. That is, if the predictions generated by the analysis lead to good decisions. If students were interested in learning about paradata analysis, then I would suggest that they gain exposure to methods in statistics, machine learning, operations research, and an emerging category "data science." It seems ...

WSS Mini-Conference on Paradata

Next week, after the big storm, the Washington Statistical Society is sponsoring a mini-conference: "Benefits and Challenges in Using Paradata." The program is available online. This will be a nice opportunity to meet and discuss with folks working on similar problems. We are few in number. It's good to take advantage of these opportunities. I'm going to be speaking about problems with working with incoming streams of paradata. I can propose some solutions, but we need to get better at this.

"Call Scheduling Algorithms" = Call Scheduling Algoritms + Staffing

We think about call scheduling algorithms as a set of rules about when cases should be called. However, staffing is the other half of the problem. For the rules to be implemented, the staff making the calls need to be there. And, there can also be issues if the staff is too large. The rules need to account for both of these situations. Probably the more difficult problem is a staff that is too large. For example, imagine that all active cases have been called. There is an appointment in 45 minutes. The interviewer can wait, or call cases that have already been called on this shift. Calling cases again would be inefficient and a violation of a rule of the algorithm. Still, it seems bad to not make he calls. I wrote a paper on a call scheduling algorithm. I assigned a preferred calling window to each case. These windows changed over time as calls were placed and the results of previous calls were used to inform the assignment of preferred window. I spent a lot of time analyzing data ...

Responsive Design and Surveys with Short Time Frames

Another interesting question that I had during the webinar that I recently gave concerned responsive design and surveys with short time frames. I have to say, I mostly work on surveys with relatively long time frames. The shortest data collection that I have worked on in the last few years is about one month. That's not to say that I think responsive design is not relevant for surveys with short field periods. I think it is. If anything, following the prescribed regimen may be more important. A key aspect of responsive design, in my mind, is that the process is pre-planned. The indicators that are monitored, the decision rules for implementing interventions, the interventions, all have to be pre-planned. In a short survey, this is particularly important as their isn't time for developing ad hoc solutions. In a former life, I worked on surveys that had field periods of a day or two. In those studies, there wouldn't have been time to meet, discuss, and decide. Given the s...

Responsive Design Definition

I've been getting ready to give a webinar on responsive design. I enjoy getting ready for this kind of talk as it gives me an opportunity to think about definitions and concepts. A few years ago, Mick Couper and I had a paper on "responsive vs adaptive" design. My thinking hasn't evolved much since that paper. In preparing the talk, I thought it might be helpful to define responsive design by contrast with ... that which is not responsive design. The contrasts were 1) pre-specified designs, and 2) ad hoc designs. The first category is a design where a pre-specified design is implemented and the results are pretty much as predicted. I personally haven't worked on many surveys like that, but I'm not yet ready to call it a "straw man." The second category is an approach I have seen in action. I sometimes call this approach "shooting from the hip." This is the situation where we start with a pre-specified design, but when it goes off the...

What is Current Standard Practice for Surveys?

In clinical trials, they have the concept that there is an "existing standard of care." New treatments are compared experimentally to this treatment. I suppose that clinical trials have some issues where informed persons can disagree about the existing standard of care, but there is at least some consensus. I'm wondering what we have for existing standard practice in the administration of surveys? As I think about running experiments, the contrast is usually to the other thing we would normally do. But, that can be ill-defined. For instance, when running experiments in our telephone facility, it was difficult to describe current practice precisely as it involved expert knowledge of the managers adjusting parameters of the calling algorithm. As further evidence that it's difficult to precisely define the essential survey conditions, there are several articles on " house effects ," where the same survey with the same (rough?) specification ends up getting ...

Reflecting the Uncertainty in Design Parameters

I've been thinking about responsive design and uncertainty. I know that when we teach sample design, we often treat design parameters as if they were known. For example, if I do an optimal allocation for a stratified estimate, I assume that I know the population element variances for each stratum. The same thing could be said about response rates, which relate to the expected final sample size. Many years ago, the uncertainty might have been small about many of these parameters. But responsive design became a "thing" largely because this uncertainty seemed to be growing. The question then becomes, how do we acknowledge and even incorporate this uncertainty into our designs? Especially responsive designs. It seems that the Bayesian approach is a natural fit for this kind of problem. Although I can't find a copy online, I recall a paper that Kristen Olson and Trivellore Raghunathan presented at JSM in 2005. They suggested using a Bayesian approach to update estimate...

Adaptive Designs and Incentives

I've been working on a paper about an incentive experiment that we did. It raised some interesting issues. And made me recall one of my favorite papers. Trussell and Lavrakas looked at incentives to a follow-up survey. They found that if someone had refused or been difficult to contact in the initial, screening survey, then a higher incentive was needed than for someone who had not refused or been difficult to contact. The incentives they recommend also differed by some demographic characteristics as well. I liked this example since the adaptation was linked, in part, to the paradata. These are the kinds of adaptations I have the most interest in. They require learning on the part of the survey organization that happens during data collection. I have the feeling that these kinds of adaptations can be particularly powerful since in models predicting response, it is often the case that paradata overwhelm the predictive power of demographic characteristics. There are all sorts of...

Understanding "Randomly Selected"

I had the opportunity this morning to meet with a medical researcher who runs many clinical trials. He spoke about the problems of explaining randomization when enrolling persons in a trial. It's hard to be sure they understand the concept of randomization. To be sure, it's even more difficult to be sure they understand the consequences of either enrolling or not enrolling in a trial. But the problem of explaining randomization caught my attention. This reminds me of the situation that interviewers find themselves in quite frequently. In implementing random selection of a person from within a household, they often find that the person selected is someone other than the informant who aided with the selection. In these cases, the informant may be disappointed that they weren't selected and ask if they can do the interview instead. It's often difficult to explain why we want to speak to the other person, who is not there or maybe not even willing to do the interview. I...

Tiny Data...

I came across this interesting po st about building a Bayesian model with careful specification of priors. The problem is that they have "tiny" data. So the priors play an important role in the analysis. I liked this idea of "tiny" data. The rush to solve problems for "big data" has obscured the fact that are interesting problems for situations where you don't have much data. Frost Hubbard and I looked at a related problem in a recently published article . We look at the problem of estimating response propensities during data collection. In the early part of the data collection, we don't have much data to estimate these models. As a result, we would like to use "prior" data from another study. However, this prior information needs to be well-matched to the current study -- i.e. have the same design features, at least approximately. This doesn't always work. For example, I might have a new study with a different incentive than I...

Interviewer Travel and New Forms of Data

The Director of the Census Bureau, John Thompson, recently blogged about a field test for the 2020 Decennial Census Nonresponse Follow-up. They are testing a number of new features, including the use of smartphones in data collection. I've been working with GPS data from smartphones used by field interviewers. The data are complex, but may offer new insights into interviewer travel. Think of travel as a broad concept -- it's not just an expense or efficiency issue. The order in which calls are made may also relate to field outcomes like contact and response rates. Perhaps these GPS data can help us understand how interviewers currently make decisions about how to work their sample. For example, do they move past sampled housing units when they first arrive to the area segment? Is this action associated with higher contact rates? Of course, travel is also an expense or efficiency issue. I wouldn't want pushing for more efficient travel to interfere with other aspects o...

Decision Support and Interviewer Compliance

When I was working on my dissertation, I got interested in a field of research known as decision support. They use technical systems to help people make decisions. These technical systems help to implement complex algorithms (i.e. complicated if... then decision rules) and may include real-time data analysis. One of the reasons I got interested in this area was because I was wondering about implementing complicated decision algorithms (e.g. highly tailored, including to incoming paradata) in the field. One of the problems associated with decision support has to do with compliance. Fortunately, Kawamoto and colleagues did a nifty systematic review of the literature to see what factors were related to compliance in a clinical setting. These features might be useful in a survey setting as well. They are: 1. The decision support should be part of the workflow. 2. It should deliver recommendations not just information. 3. The support should be delivered at the time the decision is mad...

Training Works... Until it Doesn't

I recently had need for several citations showing that training interviewers works. Of course, Fowler and Mangione show that training can improve interviewer performance in delivering a questionnaire. Groves and McGonagle also show that training can have an impact on cooperation rates. But then I also thought of the example from Campanelli and colleagues where experience interviewers preferred to make call attempts during the day -- when these attempts would be less successful and despite training that other times would work better. So, an interesting question, when does training work? And when does it not?

Setting an Appointment for Sampled Units... Without their Assent

Kreuter, Mercer, and Hicks have an interesting article in JSSAM. In a panel study, the Medical Expenditure Panel Survey (MEPS). They note my failed attempt to deliver recommended calling times to interviewers. They had a nifty idea... preload the best time to call as an appointment. Letters were sent to the panel members announcing the appointment. Good news. This method improved efficiency without harming response rates. There was some worry that setting appointments without consulting the panel members would turn them off, but that didn't happen. It does remind me of another failed experiment I did a few years ago. Well, there wasn't an experiment, just a design change. We decided that it would be good to leave answering machine messages on the first telephone call in an RDD sample. In the message, we promised that we would call back the next evening at a specified time. Like an appointment. Without experimental evidence, it's hard to say, but it did seem to increase...

Tracking: Does Sequence Matter?

I've wanted to run an experiment like this for a while. When we do tracking here, we either run a standard protocol. This protocol is a series of "tracking steps" that are carried out in a specific order. The other way we do this is to let the tracking team decide which order to run the steps in. In cases where we run a standard protocol, experts decide which order to run the steps in. Generally, the cheapest steps are first on the list. The problem is that you can't evaluate the effectiveness of each step because they all deal with different subgroups (i.e. those that didn't get found on the previous step). I only know of one experiment that varied the order of steps. Well, I finally found one that wasn't too objectionable. I got them to vary the order. We recently finished the survey and found that... the original order worked better. The glass half full view: it did make a difference which order you used. And the experts did choose that one.

Defining phases

I have been working on a presentation on two-phase sampling. I went back to an old example from an RDD CATI survey we did several years ago. In that survey, we defined phase 1 using effort level. The first 8 calls were phase 1. A subsample of cases was selected to receive 9+ calls. It was nice in that it was easy to define the phase boundary. And that meant that it was easy to program. But, the efficiency of the phased approach relied upon their being differences in costs across the phases. Which, in this case, means that we assume that cases in phase two require similar levels of effort to be completed. This is like assuming a propensity model with calls as the only predictor. Of course, we usually have more data than that. We probably could create more homogeneity in phase 2 by using additional information to estimate response probabilities. I saw Andy Peytchev give a presentation where they implemented this idea. Even just the paradata would help. As an example, consider two...

What would a randomized call timing experiment look like?

It's one thing to compare different call scheduling algorithms. You can compare two algorithms and measure the performance using whatever metrics you want to compare (efficiency, response rate, survey outcome variables). But what about comparing estimated contact propensities? There is an assumption often employed that these calls are randomly placed. This assumption allows us to predict what would happen under a diverse set of strategies -- e.g. placing calls at different times. Still, this had me wondering what a really randomized experiment would look like. The experiment would be best randomized sequentially as this can result in more efficient allocation. We'd then want to randomize each "important" aspect of the next treatment. This is where it gets messy. Here are two of these features: 1. Timing. The question is, how to define this. We can define it using "call windows." But even the creation of these windows requires assumptions... and tradeo...