Skip to main content

Posts

Showing posts with the label Experiment

Mechanisms of Mode Choice

Following up yet again, on posts about how people choose modes. In particular, it does seem that different subgroups are likely to respond to different modes at different rates. Of course, with the caveat that it's obviously not just the mode, but also how you get there that matters. We do have some evidence about subgroups that are likely to choose a mode. Haan, Ongena, and Aarts examine an experiment where respondents to a survey are given a choice of modes. They found that full-time workers and young adults were more likely to choose web over face-to-face. The situation is an experimental one that might not be very similar to many surveys: Face-to-face and telephone recruitment to the choice of face-to-face or web survey. But at least the design allows them to look at who might make different choices. It would be good to have more data on persons making the choice in order to better understand the choice. For example, information about how much they use the internet might ...

Is there such a thing as "mode"?

Ok. The title is a provocative question. But it's one that I've been thinking about recently. A few years ago, I was working on a lit review for a mixed-mode experiment that we had done. I found that the results were inconsistent on an important aspect of mixed-mode studies -- the sequence of modes. As I was puzzled about this, I went back and tried to write down more information about the design of each of the experiments that I was reviewing. I started to notice a pattern. Many mixed-mode surveys offered "more" of the first mode. For example, in a web-mail study, there might be 3 mailings with the mail survey and one mailed request for a web survey. This led me to think of "dosage" as an important attribute of mixed-mode surveys. I'm starting to think there is much more to it than that. The context matters  a lot -- the dosage of the mode, what it may require to complete that mode, the survey population, etc. All of these things matter. Still, we...

Should exceptions be allowed in survey protocol implementation?

I used to work on a CATI system (DOS-based) that allowed supervisors to release cases for calling through an override mechanism. That is, the calling algorithm had certain rules that kept cases out of the calling queue at certain times. The main thing was if something had been called and was a "ring-no-answer," then the system wouldn't allow it to be called (i.e. placed in the calling queue) until 4 hours had passed. But supervisors could override this and release cases for calling on a case-by-case basis. This was handy -- when sample ran out, supervisors could release more cases that didn't fall within the calling parameters. This kept interviewers busy dialing. Recently, I've started to think about the other side of such practices. That is, it is more difficult to specify the protocol that should be applied when these exceptions are allowed. Obviously, if the protocol is not calling a case less than four hours after a ring-no-answer, then the software explicit...

Methodology on the Margins

I'm thinking again about experiments that we run. Yes, they are usually messy. In my last post, I talked about the inherent messiness of survey experiments that is due to the fact that surveys have many design features to consider. And these features may interact in ways that mean we can't simply pull out an experiment on a single feature and generalize the result to other surveys. But I started thinking about other problems we have with experiments. I think another big issue is that methodological experiments are often run as "add-ons" to larger surveys. It's hard to obtain funding to run a survey just to do a methodological experiment. So, we add our experiments to existing surveys. The problem is that this approach usually creates a limitation. The experiments can't risk creating a problem for the survey. In other words, they can't lead to reductions in response rates or threaten other targets that are associated with the main (i.e. non-methodologic...

Messy Experiments

I have this feeling that survey experiments are often very messy. Maybe it's just in comparison to the ideal type -- a laboratory with a completely controlled environment where only one variable is altered between two randomly assigned groups. But still, surveys have a very complicated structure. We often call this the "essential survey conditions." But that glib phrase might hide some important details. My concern is that when we focus on a single feature of a survey design, e.g. incentives, we might come to the wrong conclusion if we don't consider how that feature interacts with other design features. This matters when we attempt to generalize from published research to another situation. If we only focus on a single feature, we might come to the wrong conclusion. Take the well-known result -- incentives work! Except that the impact of incentives seems to be different for interviewer-administered surveys than for self-administered surveys. The other features of...

Balancing Response through Reduced Response Rates

A case can be made that balanced response -- that is, achieving similar response rates across all the subgroups that can be defined using sampling frame and paradata -- will improve the quality of survey data. A paper that I was co-author on used simulation with real survey data to show that actions that improved the balance of response usually led to reduced bias in adjusted estimates. I believe the case is an empirical one. We need more studies to speak more generally about how and when this might be true. On the other hand, I worry that studies that seek balance by reducing response rates (for high-responding groups) might create some issues. I see two types of problems. First, low response rates are generally easier to achieve. It takes skills and effort to achieve high response rates. The ability to obtain high response rates, like any muscle, might be lost if it is not used. Second, if these studies justify the lower response rate by saying that estimates are not significantly ...

Timing of the Mode Switch

I just got back from JSM where I presented the results of an experiment that varied the timing of the mode switch in a web-telephone survey. I'm not going to talk about the results of the experiment in this post, just the premise. The concern that motivated the experiment had to do with the possibility that longer delays before switching modes could have adverse effects on response rates. This could happen for several reasons. If there is pre-notification, then the effect of the prenote on response to the second mode might be reduced with longer delays before switching.  If the first mode is annoying in some way, it can diminish the effectiveness of the second mode. The latter case is particularly interesting to me. It points to the ways that different treatment sequences can have different levels of effectiveness. We saw an impact like this in an experiment we did of two sequences of modes for a screening survey. The two sequences functioned about the same in terms of respo...

What is Current Standard Practice for Surveys?

In clinical trials, they have the concept that there is an "existing standard of care." New treatments are compared experimentally to this treatment. I suppose that clinical trials have some issues where informed persons can disagree about the existing standard of care, but there is at least some consensus. I'm wondering what we have for existing standard practice in the administration of surveys? As I think about running experiments, the contrast is usually to the other thing we would normally do. But, that can be ill-defined. For instance, when running experiments in our telephone facility, it was difficult to describe current practice precisely as it involved expert knowledge of the managers adjusting parameters of the calling algorithm. As further evidence that it's difficult to precisely define the essential survey conditions, there are several articles on " house effects ," where the same survey with the same (rough?) specification ends up getting ...

Adaptive Designs and Incentives

I've been working on a paper about an incentive experiment that we did. It raised some interesting issues. And made me recall one of my favorite papers. Trussell and Lavrakas looked at incentives to a follow-up survey. They found that if someone had refused or been difficult to contact in the initial, screening survey, then a higher incentive was needed than for someone who had not refused or been difficult to contact. The incentives they recommend also differed by some demographic characteristics as well. I liked this example since the adaptation was linked, in part, to the paradata. These are the kinds of adaptations I have the most interest in. They require learning on the part of the survey organization that happens during data collection. I have the feeling that these kinds of adaptations can be particularly powerful since in models predicting response, it is often the case that paradata overwhelm the predictive power of demographic characteristics. There are all sorts of...

Mixed-Mode Surveys: Nonresponse and Measurement Errors

I've been away from the blog for a while, but I'm back. One of the things that I did during my hiatus from the blog was to read papers on mixed-mode surveys. In most of these surveys, there are nonresponse biases and measurement biases that vary across the modes. These errors are almost always confounded. An important exception is Olson's paper . In that paper, she had gold standard data that allowed her to look at both error sources. Absent those gold standard data, there are limits on what can be done. I read a number of interesting papers, but my main conclusion was that we need to make some assumptions in order to motivate any analysis. For example, one approach is to build nonresponse adjustments for each of the modes, and then argue that any differences remaining are measurement biases. Without such an assumption, not much can be said about either error source. Experimental designs certainly strengthen these assumptions, but do not completely unconfound the sources ...

Happy Halloween!

OK. This actually a survey-related post. I read this short article about an experiment where some kids got a candy bar and other kids got a candy bar and a piece of gum. The latter group was less happy. Seems counter-intuitive, but in the latter group, the "trajectory" of the qaulity of treats is getting worse. Turns out that this is a phenomenon that other psychologists have studied. This might be a potential mechanism to explain why sequence matters in some mixed-mode studies. Assuming that other factors aren't confounding the issue.

Quantity becomes Quality

A big question facing our field is whether it is better to adjust data collection or do post-data collection adjustments to the data in order to reduce nonresponse bias. I blogged about this a few months ago. In my view, we need to do both. I'm not sure how the argument goes that says we only need to adjust at the end. I'd like to hear more of that. In my mind, it must be an assumption that once you condition on the frame data, the biases disappear and that assumption is valid at all points during the data collection. That must be a caricature -- which is why I'd like to hear more of the argument from a proponent of the view. In my mind, that assumption may or may not be true. That's an empirical question. But it seems likely that at some point in the process of collecting data, particularly early on, that assumption is not true. That is, the data are NMAR, even when I condition on all my covariates (sampling frame and paradata). Put another way, in a cell adjustmen...

Decision Support and Interviewer Compliance

When I was working on my dissertation, I got interested in a field of research known as decision support. They use technical systems to help people make decisions. These technical systems help to implement complex algorithms (i.e. complicated if... then decision rules) and may include real-time data analysis. One of the reasons I got interested in this area was because I was wondering about implementing complicated decision algorithms (e.g. highly tailored, including to incoming paradata) in the field. One of the problems associated with decision support has to do with compliance. Fortunately, Kawamoto and colleagues did a nifty systematic review of the literature to see what factors were related to compliance in a clinical setting. These features might be useful in a survey setting as well. They are: 1. The decision support should be part of the workflow. 2. It should deliver recommendations not just information. 3. The support should be delivered at the time the decision is mad...

Web Panels vs Mall Intercepts

I saw this interesting article that just came out. It called to my mind a talk that was hosted here a few (8?) years ago. The talk was someone from a major corporation who talked about how they switched product testing from church basements to online panels. They found that once they switched, the data became worse. The online panels picked products that ended up failing at higher rates. This seemed like a tough problem. There isn't much of a "nonresponse" kind of relationship here. But at least understanding the mechanism that got people into online panels and how they were then selected and agreed to participate in this kind of product testing seemed important. It's not my area, so I'm wondering if this has ever been done. Not that anyone would understand the process of recruiting people to participate in product testing in church basements. But that process at least worked. This new article looks at an old process -- mall intercepts -- for recruiting people...

Idenitfying all the components of a design, again...

In my last post I talked about identifying all the components of a design. At least identifying them is an important step if we want to consider randomizing them. Of course, it's not necessary... or even feasible... or even desirable to do a full factorial design for every experiment. But it is still good to at least mentally list the potentially active components. I first started thinking about this when I was doing a literature review for a paper on mixed mode designs. Most of these designs seemed to confound some elements of the design. The main thing I was looking for -- could I find any examples where someone had just varied the sequence of modes? The problem was that most people also varied the dosage of modes. For example, in a mixed mode web-telephone design, I could find studies that had web-telephone and telephone-web comparisons, but these sequences also varied the dosage. So, telephone first gets up to 10 calls, but telephone second gets 2 calls. Web first gets 3 emai...

Identifying all the active components of the design...

I've been reading papers on email prenotification and reminders. They are very interesting. There are usually several important features for these emails: how many are sent, the lag between messages, the subject line, the content of the email (length etc.), the placement of the URL, etc. A full factorial design with all these factors is nearly impossible. So folks do the best they can and focus on a few of these features. I've been looking at papers on how many messages were sent, but I find that the lag time between message also varies a lot. It's hard to know which of these dimensions is the "active" component. It could be either, both, and may even be synergies (aka "interactions") between the two (and between other dimensions of the design as well). Linda Collins and colleagues talk about methods for identifying the "active components" of the treatments in these complex situations. Given the complexity of these designs, with a large nu...

"Go Big, or Go Home."

I just got back from JSM, where I participated in a session on adaptive design. Mick Couper served as a discussant for the session. The title of this blog post is one of the points from his talk. He said that innovative, adaptive methods need to show substantial results. Otherwise, it won't be convincing. As he pointed out, part of the problem is that we are often tinkering with marginal changes on existing surveys. These kinds of changes need to be low risk, that is, they can't cause damage to the results and should only help. However, these kinds of changes are often limited in what they can do. His point was to make some big changes that will show big effects may require some risk. This made sense to me. It would be nice to have some methodological studies that aren't constrained by the needs of an existing survey. I suppose this could be a separate, large sample with the same content as an existing survey. However, I wonder if this is a chicken or egg type of problem....

Classification Problems with Daily Estimates of Propensity Models

A few years ago, I ran several experiments with a new call-scheduling algorithm. You can read about it here . I had to classify cases based upon which call window would be the best one for contacting them. I had four call windows. I ranked them in order, for each sampled number, from best to worst probability of contact. The model was estimated using data from prior waves of the survey (cross-sectional samples) and the current data. For a paper that will be coming out soon, I looked at how often these classifications changed when you used the final data compared to the interim data. The following table shows the difference in the two rankings: Change In Ranking Percent 0 84.5 1 14.1 2 1.4 3 0.1 It looks like the rankings didn't change much. 85% were the same. 14% changed one rank. What is difficult to know is what difference these classification errors might make in the o...

"Failed" Experiments

I ran an experiment a few years ago that failed. I mentioned it in my last blog post. I reported on it in a chapter in the book on paradata that Frauke edited. For the experiment, I offered a recommended call time to interviewers. The recommendations were delivered for a random half of each interviewer's sample. They followed the recommendations at about the same rate whether they saw them or not (20% compliance). So, basically, they didn't follow the recommendations. In debriefings, interviewers said "we call every case every time, so the recommendations at the housing unit were a waste of time." This made sense, but it also raised more questions for me. My first question was, why don't the call records show that? Either they exaggerated when they said they call "every" case every time. Or, there is underreporting of calls. Or both. At that point, using GPS data seemed like a good when to investigate this question. Once we started examining the GP...

Tracking, Again

Last week, I mentioned an experiment that we ran with changing the order of tracking steps. I noted that the overall result was that the original, expert-chosen order worked better than the new, proposed order. In this example, the costs weren't all that different. But I could imagine situations where there are big differences in the costs between the different steps. In that case, the order could have big cost implications. I'm also thinking that a common situation is where you have lots of cheap (and somewhat ineffective steps) and one expensive (and effective) step. I'm wondering if it would be possible to identify cases that should skip the cheap treatments and go right to the expensive treatment. Just as a cost savings measure. It would have to result in the same chance of locating the person. In other words, the skipped steps would have to have the same or less information than the costly step. My hunch is that such situations actually exist. The trick is finding ...