0% found this document useful (0 votes)
10 views16 pages

Abstinence Education Evaluation Insights

The article summarizes an evaluation of four abstinence education programs that received federal funding. The evaluation was conducted over 10 years with a rigorous experimental design. It found that the programs did not increase rates of sexual abstinence or reduce risky sexual behavior among youth participants. While supporters dismissed the findings, opponents cited it in debates around continued funding for abstinence education programs. The evaluation received awards for its high quality methods and stakeholder involvement within a politically charged environment.

Uploaded by

Azan Nia
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views16 pages

Abstinence Education Evaluation Insights

The article summarizes an evaluation of four abstinence education programs that received federal funding. The evaluation was conducted over 10 years with a rigorous experimental design. It found that the programs did not increase rates of sexual abstinence or reduce risky sexual behavior among youth participants. While supporters dismissed the findings, opponents cited it in debates around continued funding for abstinence education programs. The evaluation received awards for its high quality methods and stakeholder involvement within a politically charged environment.

Uploaded by

Azan Nia
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

American Journal of Evaluation [Link]

com/

Evaluation Exemplar: The Critical Importance of Stakeholder Relations in a National, Experimental Abstinence Education Evaluation
Paul R. Brandon, Nick L. Smith, Christopher Trenholm and Barbara Devaney American Journal of Evaluation 2010 31: 517 originally published online 9 September 2010 DOI: 10.1177/1098214010382769 The online version of this article can be found at: [Link]

Published by:
[Link]

On behalf of:

American Evaluation Association

Additional services and information for American Journal of Evaluation can be found at: Email Alerts: [Link] Subscriptions: [Link] Reprints: [Link] Permissions: [Link] Citations: [Link]

>> Version of Record - Oct 26, 2010 Proof - Sep 9, 2010 What is This?

Downloaded from [Link] by azan nia on November 21, 2011

Exemplars

Evaluation Exemplar: The Critical Importance of Stakeholder Relations in a National, Experimental Abstinence Education Evaluation
Paul R. Brandon1, Nick L. Smith2, Christopher Trenholm3, and Barbara Devaney3

American Journal of Evaluation 31(4) 517-531 The Author(s) 2010 Reprints and permission: [Link]/[Link] DOI: 10.1177/1098214010382769 [Link]

Evaluation studies are often considered exemplary because of the quality of their designs, but it is well known that a commonand often necessaryaspect of outstanding studies is stakeholder involvement. Furthermore, the implementation of designs and the nature of stakeholder involvement depend in large part on political and ideological contexts. The evaluation that we describe and discuss here is an exemplar of a study with a high-quality design and intensive stakeholder involvement within a highly charged political environment. The study examined sexual abstinence education, an approach that has been widely discussed in many media for about the past 10 years. The studys methods were widely lauded, and its results were accepted by stakeholders across the ideological spectrum. In this Exemplars section entry, we describe the evaluation and the evaluated programs, provide additional material about the study in the form of an interview of the evaluations two lead authors (Trenholm and Devaney) by the section editors (Brandon and Smith), and provide our collaborative reflections on the study.

Introduction and Background to the Evaluation


Beginning in Fiscal Year 1998, Title V, Section 510 of the Personal Responsibility and Work Opportunity Reconciliation Act of 1996 allocated $50 million annually in federal funding for programs that teach abstinence from sexual activity outside of marriage. The funding was originally authorized from 1998 to 2002 in the form of optional block grants to the states and has been renewed twice, most recently as part of the health reform legislation recently signed into law by President
1 2

University of Hawaii at Mnoa, Honolulu, HI a Syracuse University, NY 3 Mathematical Policy Research, Princeton, NJ Corresponding Author: Paul R. Brandon, 1776 University Avenue, Honolulu, HI 96822, USA Email: brandon@[Link]

Downloaded from [Link] by azan nia on November 21, 2011

517

518

American Journal of Evaluation 31(4)

Obama. The recent authorization will provide optional block-grant funding to states for an additional 5 years, from 2010 to 2014. A team from Mathematica Policy Research (Trenholm et al., 2007) completed an experimentally based impact study of four promising abstinence programs funded through the original Title V authorization. The study estimated the effects of the programs on sexual abstinence, risks of pregnancy and of sexually transmitted diseases (STDs), and other outcomes among participating youth. The studys findings showed no evidence that the programs had increased rates of sexual abstinence or otherwise reduced sexual risk, nor evidence that they had reduced rates of contraception, a concern that had often been cited by Title V opponents. The findings were released near the time that the Title V funding was up for its second reauthorization and became a major source of evidence in the ongoing debate over the effectiveness of abstinence programs. Many supporters of abstinence-until-marriage programs dismissed the findings, characterizing them as isolated and out of date, while opponents continue to use them as a basis for criticizing the continued funding of abstinence-until-marriage programs. The study received the American Evaluation Associations (AEA, 2009) Outstanding Evaluation Award. The announcement of the award (AEA, 2009) included the statement,
The study has been widely cited in news articles and opinion pieces in the Washington Post, USA Today, NPR, BBC and ABC News. Newsweeks Sharon Begley in an article entitled Just Say Noto Bad Science [Begley, 2007] pointed to the evaluation as a model of good research in a field where it had been sorely lacking.

The 2007 study was part of a larger, Congressionally mandated evaluation of Title V abstinence education programs conducted by Mathematica and the University of Pennsylvania (UPenn) on behalf of the Assistant Secretary for Planning and Evaluation within the Department of Health and Human Services (DHHS). The evaluation was funded through a competitive contracting process in 1998 and ultimately continued for 10 years, at a total cost of $7.6 million. The Mathematica and UPenn evaluation teams were led, respectively, by Barbara Devaney and Rebecca Maynard, who together brought vast experience designing and conducting rigorous impact evaluations of programs targeting at-risk populations. In addition to DHHS, audiences for the evaluation included federal and state policy makers, program stakeholders, school districts and other local leadership, researchers, and the general public.

Description of the Abstinence Education Programs1


The Abstinence Education legislation required that programs must teach 1. 2. 3. 4. 5. 6. 7. 8.
518

the social, psychological, and health gains from abstaining from sexual activity. that abstinence outside marriage is the expected standard for school-age children. that abstinence from sexual activity is the only certain way to avoid out-of-wedlock pregnancy, STDs, and other associated health problems. that monogamous relationships among faithful married couples are the expected standard of sexual activity. that sexual activity outside the context of marriage is likely to have harmful psychological and physical effects. that bearing children out of wedlock is likely to have harmful consequences for the child, the childs parents, and society. how to reject sexual advances and how alcohol and drug use increases vulnerability to sexual advances. the importance of attaining self-sufficiency before engaging in sexual activity.
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

519

Programs were not allowed to promote or encourage contraceptive use. The underlying rationale was that programs that encourage both abstinence and contraceptives send a mixed signal to teens; promoting the use of contraceptives was thought to undermine the abstinence message and might actually result in increased sexual behavior. The evaluation examined the effects of four abstinence education programs that addressed the eight characteristics: (a) My Choice, My Future!, in Powhatan, Virginia; (b) ReCapturing the Vision, in Miami, Florida; (c) Families United to Prevent Teen Pregnancy, in Milwaukee, Wisconsin; and (d) Teens in Control, in Clarksdale, Mississippi. These programs offered a range of implementation settings and strategies that reflected the variety of Title V, Section 510 programs nationwide. (A previous evaluation report described the implementation of these and other Title V-funded programs; see Devaney, Johnson, Maynard, & Trenholm, 2002.) The programs can be characterized in various ways: 1. 2. They differed substantially in their settings, program types, and attendance requirements. In three of the communities, the youth served were predominantly African American or Hispanic and from poor, single-parent households; in the fourth, the youth were mostly White, non-Hispanic youth from working- and middle-class two-parent households. One program served youth after school; the other three programs served youth in classrooms during the school day. Two programs were elective and two were not, leading to differences in participation and attendance levels. (Since many abstinence programs are voluntary, nonparticipation is likely to be an issue; thus, it was important to include voluntary programs.) The four programs in the evaluation had different levels of service dosage, with ReCapturing the Vision and Families United to Prevent Teen Pregnancy offering the highest service dosage and My Choice, My Future! and Teens in Control offering lower service dosage. In the Families United to Prevent Teen Pregnancy program, 43% of youth assigned to the program group did not participate in any classes, and in ReCapturing the Vision, 35% of youth assigned to the program group did not participate. My Choice, My Future! and Teens in Control were mandatory programs and youth were required to attend. Except for school absences and other minor class absence, dosage did not vary significantly across students within these two programs. Two programs served children in the upper elementary grades and two served middle schoolers. All four programs offered more than 50 contact hours, making them relatively intense among programs funded by Title V, Section 510 grants; two met every day of the school year, with one serving children participating for as long as 4 years. Two were in communities with a wide variety of health, family life, and sex education services available through the public schools; the other two programs were in schools with limited services. The programs shared many emphases such as teaching about physical development, reproduction, awareness and avoidance of risks, goal setting, decision making, and developing interpersonal skills.

3. 4.

5. 6.

7.

8.

The experimental impact design was a federal requirement for study sites, so only sites willing and able to support a random assignment evaluation were selected for the impact evaluation. Many other promising programs existed, especially some interesting community saturation initiatives. Although these types of programs were included in the earlier overall implementation evaluation (Devaney et al., 2002), if random assignment was not possible, those programs could not be included in the impact evaluation. Three of the four programs were delivered in a school setting and the fourth program was delivered in an after-school setting. Two of the four programs were delivered to upper elementary school students. These factors affected several evaluation processes:
Downloaded from [Link] by azan nia on November 21, 2011

519

520 1. 2. 3.

American Journal of Evaluation 31(4) Active parental consent was required. This took considerable resources. The school (or after-school) setting allowed for efficient group-based data collection procedures. Follow-up surveys were administered in group sessions using paper-and-pencil methods. Because the programs targeted middle-school and upper elementary school youth, where the prevalence of sexual activity is fairly low, the follow-up period had to continue long enough to capture key behavioral outcomes.

Evaluation Questions, Design, and Method


The abstinence education evaluation addressed three primary questions: (a) What impacts do programs have on sexual abstinence and sexual activity? (b) What impacts do programs have on possible mediators of behavior, including knowledge of pregnancy, STD risks, effects on health, opinions about abstinence, and relations with peers? and (c) What are the links between the possible mediators and behavior? As required by Congress, the evaluators used an experimental design to address these questions, with eligible youth within each of the four programs randomly assigned to either a treatment group that was offered the abstinence education program or a control group that was not offered the abstinence education program. The evaluators prepared a logic model to assist in designing the study. A survey questionnaire was administered to 2,057 youth approximately 46 years after beginning their participation in the study. There were 1,209 (59%) students in the program group and 848 in the control group. The questionnaire addressed sexual behavior, including rates of abstinence, rates of unprotected sex, number of sexual partners, expectations to abstain, and reported rates of pregnancy, births, and STDs. It also addressed knowledge and perceptions of risks of teen sexual activity, including scales about STD identification (from a list of diseases), risks of pregnancy and STDs from unprotected sex, health consequences of STDs, and perceptions of the effectiveness of condoms and birth control pills for pregnancy prevention and for the prevention of several types of STDs. The evaluation included an extended follow-up period for the purpose of examining program effects on sexual abstinence and activity. The evaluation included four waves of data collectiona baseline and three follow-up surveys. At the time of the final follow-up survey, administered 46 years after youth enrolled in the study sample, the age of the study youth ranged from 12 to 20 years, with a mean of about 16.5 years. All the surveyed youth had completed their programs, in some cases several years earlier. Given the age distribution, substantial variation was expected in the rates of sexual abstinence and sexual activity across the study sites at the time of the final follow-up survey. The estimated impacts were based on an intent to treat experimental design. This type of design provides estimates of the average impacts of offering abstinence education services to youth assigned to the program group, regardless of whether and to what extent the youth participate in the program. Put differently, this type of design does not provide estimates of program impacts only on participating youth but on all the youth offered the opportunity to participate. For the two voluntary programs, the analysis adjusted the estimated impacts for nonparticipation among those assigned to the program group. The evaluators estimated program effects as the difference in mean values between the program and control groups, adjusted for background and mediator variables measured at baseline. They estimated the effects for all four sites combined (using the average across the four sites) and for each site individually. Terms for the interaction between program site and experimental group status were entered into each model. Three sources of information helped the evaluators ascertain what the control group received. First, they collected information on attendance in the program sessions, both to assess whether participants received the services and whether control students were allowed in (referred to as
520
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

521

crossovers). Crossovers were relatively rare, thus confirming that random assignment was implemented as designed. Second, during site visits, the evaluators gathered information on the counterfactual, attempting to determine what services the control group would receive. Finally, the followup surveys included questions about reproductive health services received. It did not appear that youth received much else in the way of abstinence education or other pregnancy prevention services. As a weak proxy measure of other abstinence education involvement, for example, the proportion of youth who reported having completed an abstinence pledge ranged from 8% to 23% across the four sites. Moreover, this upper end figure is for control youth for the Families United to Prevent Teen Pregnancy program, who averaged only age 10.3 at first follow-up and so most likely included some youth who were confused about what the pledge is and reported on it incorrectly.

Evaluation Findings Effects on Behavior


The findings from the final survey showed that youth in the program group were no more likely than those in the control group to have abstained from sex. Among those who reported having had sex, the two groups had similar numbers of sexual partners and the same mean age of initial sex. Contrary to concerns raised by some critics of the Title V, Section 510 abstinence funding, the program group youth were no more likely to have engaged in unprotected sex than the control group youth. Moreover, these results from the final wave of the follow-up survey had also been found in each previous wave of the survey, suggesting no impacts both in the short-term and long-term. As described earlier, two of the programs were voluntary and had significant levels of nonparticipation. In this situation, the typical method used to adjust impact estimates for nonparticipation is to divide the impact estimates by the participation rate (Bloom, 1984). The implication of this adjustment is that the estimates of program impacts for the subsample of program participants will be larger than the estimates for the entire sample in these two sites. However, because there is a corresponding loss of statistical power when estimating impacts for the smaller, participant-only sample, the statistical significance associated with these participant-only impacts is roughly equal to those for the full sample. As a result, these participant-only impact estimates were not statistically significant.

Effects on Knowledge of Risks Associated With Teen Sex


Overall, the programs improved identification of STDs but had no overall impact on knowledge of unprotected sex risks and the consequences of STDs. Both program and control group youth had a good understanding of the risks of pregnancy but a less clear understanding of STDs and their health consequences. The program and control group youth had similar perceptions about the effectiveness of condoms for pregnancy prevention. A number of youth in both the program and the control groups reported being unsure about the effectiveness of condoms at preventing STDs, but very few youth in either group reported being unsure about their effectiveness in preventing pregnancy. The groups perceptions of the effectiveness of birth control pills in preventing pregnancy and in preventing STDs were similar.

Discussion of the Findings


The evaluation highlights the challenges faced by programs aiming to reduce adolescent sexual activity and its consequences. Until recently, rates of teen sexual activity had been in decline for well over a decade, yet about half of all high school youth had reported having had sex, and more than one
Downloaded from [Link] by azan nia on November 21, 2011

521

522

American Journal of Evaluation 31(4)

in five had reported having had four or more partners by the time they graduated from high school. One quarter of sexually active adolescents nationwide have an STD, and many STDs are lifelong viral infections with no cure. Some policy makers and health educators have questioned whether the Title V, Section 510 programs focus on abstinence elevates STD risks. The findings from this study suggest that this is not the case, as the program group youth were no more likely to engage in unprotected sex than their control group counterparts. However, given the lack of program effects on behavior, the evaluators concluded that policy makers who seek effective methods to reduce the high rate of teen sexual activity and its negative consequences should focus on continued program development. In particular, findings from the study suggest that (a) targeting youth solely at young ages may not be sufficient and (b) efforts to sustain positive peer networks may have protective effects on youth behavior. Recognizing the continued need for addressing the problem of teen sexual risk behavior, federal policy makers have used the studys findings as a basis for expanding their investment in both research and program support for teen pregnancy prevention. Within a year after the studys release, a new, large-scale, multisite experimental impact evaluation of teen pregnancy prevention programs was funded. The evaluation design reflected some of the lessons learned from the Mathematica study by emphasizing the need to evaluate the effects of abstinence programs targeting high-school age youth. The guidance was modified under the Obama administration to place relatively more emphasis on evaluating comprehensive sex education programs, while continuing to recognize a need to examine the effects of the latest generation of abstinence programs through rigorous evaluation. This recognition seems entirely warranted, as many communities around the United States continue to hotly debate the value of abstinence-until-marriage programs, while others remain strongly supportive of these programs but lack thorough information on how best they can be implemented.

Interview
In written exchanges and a conference call, the Exemplars editors and the two lead evaluators gathered and consolidated additional material, presented here in an interview format that elaborates on the exemplary features of the abstinence evaluation in a manner consistent with the journal sections editorial statement (Brandon & Smith, 2010).

Technical Issues
Editors: Your description of the abstinence education evaluation focuses primarily on technical features of the four abstinence projects and of the evaluation study. This is a typical approach to reporting evaluations, but it addresses only part of what is exemplary about your study. Much more of the exemplariness has to do with contextual issues such as how you involved stakeholders, the politics surrounding the study, the use of the findings, and so forth. To do a study of this complexity that lasts 10 years and survive itwhile keeping all the parties involved, committed, and assured of your impartialitywas exemplary in itself. So in our discussion with you, we want to address both the technical aspects of the study, which will interest readers seeking to know about conducting a randomized experiment, and also about the contextual aspects, which probably will interest all evaluators. Evaluators: Thats the focus that we would like to see as well. We believe that our study was technically sound, but it was nothing exceptional relative to what our colleagues at Mathematica and other researchers do in multiyear national studies all the time. We were able to retain the research integrity during the study while at the same time strengthening our standing with the abstinence education community, including those who were disappointed with the results. Editors: Good; were on the same track here.
522
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

523

Lets start with some questions about technical features of the study that arose when we reviewed the full report and the description that we are presenting here. First, a general question about the findings: How strong is the claim attributing effects to participation in the program? Or put a different way, what rival explanations might be made? Evaluators: Given the random assignment design and high response rates at all follow-up data collection points, we are confident that the basic study results of no significant program impacts on sexual abstinence is correct. The random assignment resulted in program and control groups that were similar at baseline. The final follow-up data collection had an overall response rate of 82%, with virtually no difference in the rates between groups (83% for the program group and 82% for the control group). One of the criticisms of the evaluation was that we measured impacts 4 to 6 years after study enrollment, long after the point when youth had received program services. We disagree with this criticism because it ignores two facts. First, like the vast majority of the Title V, Section 510 programs at the time of site selection, the four programs we studied delivered services to youth during their middle-school or upper elementary school years. However, it is the high school years when many youth make decisions about sexual activity and the evaluation had to include a follow-up data collection plan of sufficient length to capture these behavioral outcomes. Second, the impact evaluation survey was the last of three follow-up surveys that we administered; the earlier follow-up surveys likewise showed no effects on behavior, but those results could not be considered conclusive because the relatively young age of the study sample meant that most would report being abstinent even without the programs. Editors: What other potential threats might there have been to the validity of the findings of the study? Evaluators: During the course of the evaluation, we considered several potential threats to the study and, when possible, took steps to mitigate them. The first potential threat had to do with the confidentiality of the survey responses. We took four steps to address the potential concern of the students about confidentiality. First, we had independent, trained professional interviewers conducting all survey data collection. No program staff members were allowed in the classroom or were allowed to see any survey questionnaires. Second, the students completed the surveys by themselves in the presence of trained and independent interviewers. Third, we asked for no personal identifying information on the survey instruments. And finally, we removed all the completed surveys from the schools as soon the students had filled them out. The second potential threat had to do with social desirability effects. We were concerned that children in the program group might feel compelled to report remaining abstinent. However, the final follow-up data collection occurred long after the children participated in the program services and at ages when a substantial proportion of youth were sexually active, which we believed would substantially mitigate social desirability effects. Third, we took steps to address potential sample attrition. Attrition could have been a threat to quality and validity because the data collection took place over so many years. The issue was even more pronounced because youth were spread over large numbers of schools and because many had moved. To address this issue, we identified which schools in a program site had significant numbers of sample members. Then we sent trained interviewers to interview students in group sessions at those schools. Finally, we moved the interviewing effort to Mathematicas phone center and embarked on an intensive telephone follow-up effort. As a result, we had remarkably high response rates of 82% at the time of the final follow-up. When we tested whether the mode of follow-up (telephone vs. field) had any effect on student responses, we found none. We also analyzed nonresponse and calculated sample weights to adjust for nonresponse in the final data analysis. The fourth threat was the potential of contamination of both the program and control groups. In all but one site (in Virginia), the likelihood of this was remote because the programs reflected a small
Downloaded from [Link] by azan nia on November 21, 2011

523

524

American Journal of Evaluation 31(4)

fraction of youth in each school. (At the Virginia site, the intervention basically split the school population down the middle between the two groups.) So, overall, the dispersal of students over large numbers of schools and locations at the times of follow-up data collection suggested that contamination would not explain the study findings. On a broader level, you might ask what the kids were doing in the control group at the time that youth in the program group were participating. (This is the literal counterfactual.) In each site, the answer to this question was not much related to sex or abstinence education. Rather, they were generally doing other activities entirely. You might also ask about what other types of services were available to both groups outside of the program. Not surprisingly, the answer here was complex. We could only get a partial handle on it through our site visits and data collection. In general it appeared that the kids received little else in the way of abstinence education. In only one site (Miami) did most kids receive something else formally (a comprehensive sex education class in the ninth grade). Thus, with this possible exception, the environment does not appear to have been at all saturated with other competing or conflicting services. Editors: Were the four sites so different as to preclude generalizable findings with respect to abstinence education, or were they somehow representative of sites nationally? Evaluators: The results of this evaluation were not intended to be generalizable to all abstinence education programs nationwide or even those funded as part of the Title V authorization. The evaluation started after a major new source of funding was authorized in 1998, and program funding varied considerably across and within states. Within 2 years of program authorization, more than 700 programs operated nationwide. Subsequent funding for abstinence-untilmarriage programs through Community Based Abstinence Education increased the number of such programs even more. The four sites in the evaluation were selected because they were considered well-implemented, intensive, and willing and able to comply with the rigorous evaluation design. During the siteselection phase of the evaluation, the project team first called and met with numerous state officials and experts across the country to identify promising programs for the evaluation. Grant applications and program documents then provided additional detail on program goals, target population, program activities, size, and curricula. The evaluation team visited and observed 28 abstinence education programs across the nation. After extensive communication with abstinence experts and DHHS staff, 11 programs were invited and agreed to participate in the overall evaluation. Of these 11 programs, 4 had the capacity to be part of the experimental impact evaluation. That is, they were operating at a scale to provide sample sizes with sufficient statistical power to be able to detect policy relevant impacts; they were willing and able to implement a design where individual students were randomly assigned by an outside evaluator to the program and control groups; and they were on board with all follow-up data collection conducted by trained, independent data collectors, as opposed to program or school staff. The programs were not a representative subset of all abstinence-until-marriage education programs, but we judged that they offered a rich range of program strategies and implementation settings. Editors: But the problem remains, what do these results tell us about the effectiveness of abstinence education programs at the middle-school level? Only that they didnt work in these four places? Isnt $7.6 million a lot of money to spend just to find that out? Evaluators: Most abstinence programs are small, so it would have taken a lot of programs and lots of time to make this a nationally representative study of abstinence programs. It would have been much, much more costly. In addition, this is a common approach of the federal government when they have a new policy initiativeto test it on a demonstration basis to see first if can work before going large scale to see if it works nationwide. Editors: So, overall, is it correct to say that you conclude that the study limitations and differences in treatment type, level, amount, etc., did not make it impossible to find any differences if they did exist?
524
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

525

Evaluators: Overall, the strength of the evaluation designmost notably, the experimental design and high response ratessuggests that the four programs selected for the evaluation had no impact on sexual abstinence and no impact on rates of unprotected sex. We think any study limitations would not change that basic result. That being said, as noted in the evaluation report, the evaluation findings provide no information about the effects that programs might have if they were implemented for high school youth or began at earlier ages but served youth through high school. Editors: In November of 2009, an independent panel appointed by the Centers for Disease and Control and Prevention reviewed 83 sex education studies, including 21 abstinence-only programs conducted between 1980 and 2007 (Guide to Community Preventive Services, 2010). They concluded that there was not enough evidence to determine the effectiveness of abstinenceonly programs in birth control and protection against disease. Was your study included in that review? Evaluators:Yes. In fact, the review broke our sites out individually, so they represented 4 of the 21. Editors: If none of the prior studies have shown the effectiveness of abstinence-only programs, how many studies, or what kind of studies, will it take to do so? Evaluators: First, many of the studies in that CDC-supported review were not experiments and so provide less convincing evidence of whether programs were effective. And there is a contrasting paper by the Heritage Foundation that cites studies that do find effects for abstinence programs, albeit again mostly studies without a randomized experimental design. So, in total, there is less evidence to draw on for assessing the effectiveness of abstinence programs than is probably recognized. There is a real need for more experimental evaluations of these and other models of teen pregnancy prevention programs to build the evidence base further. More to the point of your question, there is no magic number we can point to that can indicate definitively whether a program approach works or not. In part, this is because a well-designed and well-implemented impact evaluation of one or more abstinence programs (or comprehensive sex education programs or other models) can only draw internally valid conclusions. It cannot draw definitive conclusions about the effectiveness of other, let alone all, programs that were not a focus of the study. And, coupled with this, the programs are very diversediffering in their target populations, duration, content, settings, and other featuresand are evolving over time both in an effort to improve and to reach new or different groups of youth. So it is not really possible to evaluate rigorously all the different permutations of the programs for all possible different groups. Editors: So, if we can never amass enough information to prove that a particular approach cannot work, and we know that no single approach will ever work in all possible circumstances, what can we learn from doing national evaluations of policy alternatives? Evaluators: We can understand the desire to draw sweeping conclusions from a policy evaluation that program models work or dont work, but the reality of what we can learn is more subtle when programs are as diverse and evolving as those in the field of pregnancy prevention. In this field, research of a particular program or group of programs can largely make only internally valid statementsthat is, that the evaluated program(s) did or did not have a beneficial impact for the sample of youth under study. More generalizable findings are possible through good qualitative information and, ideally, through replication, but policy decisions based on those findings will always be open to interpretation and to (ideally honest) policy debate. For example, if a rigorous evaluation of an abstinence program finds positive effects on behavior, a good companion implementation and process study can help us understand why it may have worked and for whom we might expect it to work for in the future. And, ideally, additional studies can be conducted to measure the impacts of that same program once it has been replicated. This collective information can be immensely helpful to a policy maker trying to understand if this abstinence program is worth investing inand how it should be implemented and for whom. But it does not guarantee the success of the program in all future replications and it certainly does not
Downloaded from [Link] by azan nia on November 21, 2011

525

526

American Journal of Evaluation 31(4)

support conclusions that all abstinence programs work, let alone work better than other approaches. Conversely, an evaluation that finds no effects of a group of abstinence programs can be immensely helpful to policy makers, in part because it might be able to inform program refinement. But it likewise cannot support broad conclusions that all abstinence programs cannot work.

Context and Stakeholder Issues


Editors: Now lets discuss the context within which the evaluation occurred. Lets start with having you remind us about the national political context of the study and the focus of the abstinence education programs that you evaluated. Evaluators: Youll recall that the evaluation lasted for 10 years. It started during the Clinton administration under a Republican-controlled Congress, which was able to get large-scale funding for the abstinence education program. The funding was very controversial; there had been some abstinence programs around before, but this one was very prescriptive about what they meant by abstinence, which meant abstinence until marriage. The evaluation was not supposed to be only about teenagers, but the funding was focused on them, so they were our focus during the evaluation. Editors: Given how controversial the topic was, how were you received when you were awarded the evaluation? Evaluators: When we won the evaluation, the abstinence advocacy community was concerned that their perspectives would be ignored. In their opinion, abstinence is the message that should be given to teens and, at that point in time, there were few programs that delivered this message. They were concerned that an evaluation team had to understandperhaps even adoptthis perspective in order to do a good evaluation. Editors: How did you deal with this challenge? Evaluators: One of the main things that we did was to create and use a technical work group (TWG) that included individuals trusted by the abstinence community, as well as individuals with a broad range of policy perspectives and research expertise. For example, the group included the authors of the abstinence legislation and an obstetrician who was a very strong supportive and effective spokesperson for abstinence-only education. Some members of the TWG had done a lot of research on comprehensive sex education programs, abstinence programs, or predictors of adolescent sexual abstinence. Others were strong policy people who supported abstinence education or comprehensive sex education. We often have advisory groups on these large studies, but they are not studies that are as sensitive as this one, and the groups do not always function as well as this one did. We think the way we worked with the TWG was important, but in some ways we were lucky. It is necessary but not sufficient to involve stakeholders; we happened to have a group of stakeholders who had their positions but at the same time were reasonable and smart and willing to talk. When you have that, making sure that you constantly engage them is important. The TWG by and large got along very well. Most stayed with us through the whole study. (One key player did not.) They understood the design, and they understood the outcomes that we were measuring. We gave them briefings and we gave them preliminary results. When we released the report, we used quotations from some of them. At the end, the TWG membersincluding those who were strong supporters of abstinence educationacknowledged the quality of the study. Still, it was not easy for us or, at times, for the TWG members. In the early stages of the study, for example, several individuals who were highly respected among abstinence proponentsand who could possibly have undermined the evaluationopenly criticized our work as flawed and biased. Throughout this period, we worked diligently to address these criticismsor at least keep them from being translated to actions that would be harmful to the evaluations site recruitment or implementation. We did this mostly through direct dialogue, reaching out to these critics on a one-on-one basis
526
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

527

and trying to answer questions that they had, address any misconceptions, and help them understand that we were deeply committed to understanding these programs and to ultimately conducting highquality, objective research. We also reached out to several TWG members to speak with these critics on our behalf. Those conversations were no doubt tense at times, and we really are grateful to those TWG members for being willing to put in that kind of effort on behalf of a research study. These outreach efforts were intensive and required considerable persistence. They also required a passion for explaining how evaluation can help programs, a commitment to working with and listening to program staff and key stakeholders, and ongoing attention by very senior and seasoned staff. Rebecca Maynard was the driving force behind the outreach effort, and she was incredibly skilled at engaging programs and bringing disparate groups together for the evaluation. Editors: That is a telling description of your work with stakeholders at the national level. What about stakeholders at the local levelprogram personnel and others? Evaluators: Well, we used the same process where we identified who the real actors were at the local level and secured their cooperation and support. The way to do this is by being completely honest. You need to understand the program from the local stakeholders perspectives and why they think the program is worthy and how they implement it. And you need to get their buy-in so they believe the study is credible. Let us give you an example of getting buy-in. During the design phase of the evaluation, we worked on the logic model and what outcomes were important to measure. The prevalence of unprotected sex among sexually active youth was an important public health outcome to measure (in addition to the obvious choice of measuring whether youth remained abstinent). Often that outcome is measured by contraceptive use. The problem is that abstinence advocates do not want youth asked questions about contraceptive useespecially youth who have remained abstinent. Indeed, their argumentwhich has a lot of validityis that asking youth whether they have used contraceptives undermines the main message of abstinence education programs. On the other hand, some members of the public health community raised concerns that abstinence education places teens at risk of STDs and unwanted pregnancy by not supporting contraceptive use, so we felt it was important to include some questions on the follow-up surveys to assess whether the programs increased these risks. (They did not.) We made two important decisions that addressed the legitimate concerns expressed by abstinence supporters, as well as to satisfy the evaluation objectives, and as a result we largely succeeded in getting buy-in from the broad range of stakeholders. First, in the logic model, we listed the outcome as risks of pregnancy and STDs, as opposed to contraceptive use. This was a subtle, but important, change, as all agreed that reducing the risks of pregnancy and STDs was a good outcome, while not all agreed that contraceptive use was a good outcome. Second, in the main follow-up survey questionnaire, we asked only one question about whether the respondent had ever had sex. Depending on the confidential answer to that question, the respondents had another survey module. Students responding that they had had sex then answered questions on age of first sex, number of partners, whether they used protection, and pregnancy. Students who had not had sex responded to a different set of questions that did not ask any additional questions on sexual activity and were designed to take the same amount of time to complete as the sexual activity module. Editors: Were there any other key stakeholder groups that you dealt with? What about the evaluation client? Evaluators: The client was the HHS Assistant Secretary for Planning and Evaluation. That position heads up the research arm of the organization. We had several different project officers throughout the evaluation, and the political appointees (assistant secretaries) changed, too. The first assistant secretary was very bright, likable, articulate, and played a key role; he was beloved by both the left and right. We could not have been luckier to work with him. Later, during the Bush administration,

Downloaded from [Link] by azan nia on November 21, 2011

527

528

American Journal of Evaluation 31(4)

there were several assistant secretaries, and over time they became more conservative and some had little research background. At the time the report was done, we worked with a deputy secretary who had come from the Heritage Foundation and had written papers about abstinence education. Initially, we were nervous about our work with her because she came with such apparently strong priors about the program. But we had nothing to be nervous aboutshe believed in people, and she was trustful of us and the process. She facilitated the process throughout, including the release of the report. In fact, she worked with her staff on a very effective means for its release: A reporter from the Associated Press (AP) interviewed each one of us. He wrote the first article about the report, and then other reporters began to contact us. The key thing is that the AP story quoted us as saying that the programs did not show evidence for effectiveness in reducing abstinence levels but also did not show evidence of contraceptive use. That was very important, because a common criticism of abstinence programs is that they reduce contraception use. HHS was thrilled because both points came out and everyone was happy that we had worked our way down to a balanced press release. After that, HHS trusted us to speak on behalf of the report. Normally, the Department would take the lead. This spoke to the level of trust that HHS had in us. So there was both good fortune and we did our job well. Editors: The issues you are raising have to do with the use of evaluation findings, which we all know is a common topic in research on program evaluation. You touched on use in your discussion of how the technical work group accepted your results as well. Did you glean any other insights having to do with the use of evaluation findings? Evaluators: The supporters of abstinence programs on the TWG responded as they did in part because they had been with us from the beginning, and so it was the appropriate response for them to have. But this was not the case for all supporters of abstinence programs, and some of them, not surprisingly, were critical of the findings and questioned the quality of the work. What was interesting in hindsight, though, is that they mostly ignored the finding that the programs did not put teens at increased risk of pregnancy or STDs, due to the lack of impacts on unprotected sex. What was particularly interesting about this null finding on unprotected sexwhich was favorable for the abstinence education programsis that had we not been able to include questions on contraceptive use on the follow-up questionnaires, the evaluation would have been silent on impacts on unprotected sex. It is unclear whether the abstinence supporters ignored this result because of their concern over the finding of no impact on sexual abstinence or because they did not view the lack of impacts on unprotected sex as a good program outcome. There were a lot of dimensions to the use of the findings. For example, there was a lot said in the media and other outlets about the lack of effects of abstinence education. But they ignored that there was no effects on unprotected sex. The argument from the abstinence opponents for years had been that abstinence programs are bad because they put kids at risk. We found no evidence that this was the case. A Newsweek article said the study showed no effects on whether the kids remained abstinent but then drew on other studies showing that students in abstinence education programs use contraception less, making it look like our study had shown likewise (Quindlen, 2009). We wrote to Newsweek to refute that, but they did not publish our letter. The media used our results and then interpreted them in light of the findings of other studies that suggested abstinence education puts youth at risk.

Reflections
AEA selected the abstinence education evaluation for the associations annual best evaluation award because of the studys reputation as the best-designed and best-conducted examination of abstinence education to date. Newspapers, magazines, blogs, and other media widely reported the evaluations findings. Issues about funding sexual abstinence have fueled long-standing controversies between
528
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

529

liberal-leaning and conservative-leaning educators and policy makers, but the report was essentially accepted by those on both ends of the political spectrum. The evaluation was authorized in the late 1990s under federal legislation that mandated funding for programs adhering to a particular definition of strict abstinence. The study used a randomizedexperiment design with a substantial population of youth in four middle-school programs that the evaluators assessed as promising based on their duration, credible approach, and likely strong implementation. The programs varied by geographic regions, community settings, methods, schedules, attendance requirements, grade levels served, and curricular emphases. The study addressed the long-term effects of the programs on sexual abstinence and activity, on mediators of behavior, and on the links between the mediators and behaviors. Effects were measured in a multiwave administration of survey questionnaires that followed the students until they were old enough to engage in sexual activity. At the conclusion of the longitudinal span, the evaluators found no effects on abstinence or on the use of contraception. The evaluation first caught the attention of the Exemplars editors because of its AEA award. The technical quality of the study was immediately apparent when the editors examined the report; clearly, the study meets the requirements of an exemplary randomized study using survey methods. Furthermore, the report addresses the span of the 10-year study in clear language accessible to the layperson while satisfying the technical information requirements of researchers and evaluators. The editors interest in the study was enhanced considerably after their initial discussion with the evaluators when it became clear that the studys exemplariness was about much more than its technical adequacy. The evaluators description of the controversial nature of abstinence education and the manner in which the study unfolded makes the evaluation an obvious candidate for examining how to conduct an exemplary study under politically and ideologically contentious circumstances. So this study is exemplary, in part, for its masterful involvement of stakeholders while implementing an outstanding experimental design. The study is a contribution to the empirical research on the longstanding topic of stakeholder participation in evaluation. For about the past 15 years, much of the literature on stakeholder involvement has focused on engaging stakeholders for the purposes of enhancing the use of formative evaluation findings (mostly in local or regional studies) or of empowering the recipients of program services in their daily lives. The present account focuses on the intricacies of involving stakeholders in a highly charged political environment primarily for the purposes of enhancing the validity and utility of the evaluation and ensuring that ideological and political issues are addressed in a balanced fashion. The evaluators agreed with the editors assessment of the rationale for deeming the study exemplary. They were quick to point out to the editors that they were confident in the technical quality of the study but did not believe its technical attributes alone made the study worthy of an award any more than many large studies that they and their colleagues and others regularly conduct on a national scale. They agreed that the difficulties of addressing the concerns, viewpoints, and idiosyncrasies of ideologically and politically divergent stakeholders over the duration of a multiyear evaluative study requires exceptional insight, subtlety, hard work, and a certain degree of luck. Technical skills, professional methods, and alternative designs are learned through initial evaluation training, and descriptions of them abound. The ability to carry out studies in difficult circumstances is primarily learned through experience, however, and presenting an account of such a study can provide information useful to novices and experts alike. When dealing with powerful representatives of stakeholder groups, the stakeholder notion takes on a different connotation than it does in small formative participatory evaluations. The evaluators believe that maintaining a balance between stakeholders across the range of the political spectrum was a key aspect of making the study succeed. They recruited a technical work
Downloaded from [Link] by azan nia on November 21, 2011

529

530

American Journal of Evaluation 31(4)

group of researchers, as well as those in policy, with prior leanings for or against abstinence education. They briefed them often, shared difficulties encountered when conducting the study, and elicited their advice. They conferred with the work group about the variety of issues that evaluators typically encounter, working all the while with an added task of balancing stakeholders ideologically divergent perspectives. With the exception of one work group member who chose to disassociate himself early on from the study, they were successful in their interaction with the work group. The evaluation team also spent a considerable amount of time on site observing program operations. They understood the context of program operations and the needs and values of local stakeholders. The local and national abstinence supporters and stakeholders deserved recognition for their expertise on the programs and even acknowledgement of their common perception that the programs were misrepresented in the media and elsewhere. They also had to be managed thoughtfully. For example, the evaluators had to respect the wishes of program stakeholders not to include an explicitly stated component about contraceptive use in their logic model, because such a component was contradictory to the goals of abstinence education. At the same time, the evaluators knew that the public health community would object if no data were collected about the effects of abstinence education on the participating students sexual behavior in the long-term. The evaluators relabeled the logic model from contraceptive use to risks of pregnancy and STDs. The abstinence stakeholders concurred with this change, thereby ensuring that their perspectives were addressed while allowing the evaluators to collect data that addressed the perspectives of the public health community, as well. Labels and interpretations matter. (Ironically, the findings showed that abstinence education did not increase the risks of pregnancy and STDs, contrary to the critics of abstinence education.) The evaluators emphasized several times in their discussion with the editors that a degree of their success in balancing stakeholders needs was due to hard work, good fortune, and committed staff. They had the advantage of working with project officers and staff at the DHHS Assistant Secretary for Planning and Evaluation who were supportive of the rigor of the evaluation. Their account of how the DHHS deputy secretarys public release of the final evaluation report enhanced the studys credibility presents a clear manifestation of the quality of the government staff involved with the evaluation. Collaboration is essential at all levels; stakeholders at all levels must agree to cooperate and not to co-opt each other. The pattern of the findings of the study might be considered another fortuitous feature of the study that helped foster the wide acceptance of its credibility. The evaluators found that abstinence education had no effect on abstinence levels, to the chagrin of abstinence supporters, but neither did it affect contraceptive use, to the disappointment of opponents of abstinence education, who believed that declines in contraceptive use are an inevitable result of teaching abstinence. If the findings of only one side had prevailed, might there not have been an uproar about the results? And if so, what are the implications of this conclusion about the importance of stakeholder participation? Is close, careful, and respectful work with stakeholders likely to have such strong benefits only when all sides perceive the findings to be balanced? The answers to these questions are unclear. On a broader topic, the findings of the study are reminders of long-standing doubts about the presence or strength of the effects of many educational programs. Complaints about lack of effects of these programs have been widespread in the United States since large-scale federal programs and their evaluations were first funded in the 1960s. With the increase over about the past 10 years in the number of large randomized studies not showing effects, this state of affairs does not seem to have abated. Were the abstinence programs properly conceived? Should abstinence be taught beginning in middle schools, as was the case in the four programs that the evaluators examined, and continued until youth are most likely to begin sexual activity? Or should programs for older children be taught regularly and intensely through high school instead? Do the findings simply reflect a lack of an intervention of any degree of curricular credibility at any level?
530
Downloaded from [Link] by azan nia on November 21, 2011

Brandon et al.

531

Another aspect of no effects is the nature of the generalizability of the findings. The evaluators responses to the editors questions about this issue constitute an insightful addition to the ongoing debate about the use of randomized experiments to examine what works. The evaluators stated that the findings were not generalizable beyond the studied programs. In these circumstances, researchers and evaluators can show only the degree to which programs are effective in particular instances. Even if they are successful in some circumstances and in some settings, it cannot be shown they will always work, and when they are unsuccessful, it cannot be shown that they will never work. In the parlance of the literature on the use of evaluation findings, limits on the generalizability of the abstinence evaluation results mean that they had limited direct instrumental use, in that they do not definitively answer the what works? policy question. They did, however, have indirect instrumental use, in that they suggest which questions to investigate next such as, Would abstinence education at the high school level be effective? In addition, the findings reported had significant policy conceptual use and political influence. They contributed to extending a balanced national policy debate. It is clear that this study, exemplary in design, execution, and stakeholder involvement, has made significant contributions to the national conversation on abstinence education. Declaration of Conflicting Interests The authors declared no conflicts of interest with respect to the authorship and/or publication of this article. Funding The authors received no financial support for the research and/or authorship of this article. Note
1. The following description is adapted from the Executive Summary of the Final Report: [Link] [Link]/publications/PDFs/[Link]. The full Final Report is available at: [Link] /PDFs/[Link].

References
American Evaluation Association. (2009, November 3). AEA awards honor contributions to the field of evaluation. Retrieved from [Link] Begley, S. (2007, May 7). Just say noto bad science. Newsweek. Retrieved from [Link] 2007/05/06/[Link] Bloom, H. (1984). Accounting for no-shows in experimental designs. Evaluation Review, 8, 225-246. Brandon, P. R. & Smith, N. L. (2010). Editorial statement. American Journal of Evaluation, 31, 252-253. Devaney, B., Johnson, A., Maynard, R., & Trenholm, C. (2002). The evaluation of abstinence programs funded under title V section 510: Interim report. Princeton, NJ: Mathematica Policy Research. Guide to Community Preventive Services (2010). Prevention of HIV/AIDS, other STIs, and pregnancy. Retrieved from [Link] Quindlen, A. (2009, March 7). Lets talk about sex. Newsweek. Retrieved from [Link] 2009/03/06/[Link] Trenholm, C., Devaney, B., Fortson, K., Quay, L., Wheeler, J., & Clark, M. (2007). Impacts of four title V, section 510 abstinence education programs. Princeton, NJ: Mathematica Policy Research.

Downloaded from [Link] by azan nia on November 21, 2011

531

You might also like