Chapter 14 of 30

Continual policy evaluation —a practitioner’s guide

The Policy Playbook10 min read
47% through bookBack to overview

The Policy Playbook

Continual policy evaluation —a practitioner’s guide

Reader's guide

What this chapter gives you

A section on using evidence and evaluation to improve policy judgement, avoid false confidence and build learning into delivery.

Design evidence before implementation

Separate signal from noise

Use evaluation as a learning loop

Continual policy evaluation —a practitioner’s guide

By systematically mapping challenges, solutions, and gaps through this multi-evidence lens, public servants can make more informed decisions that honor both technical knowledge and local wisdom, scientific analysis and community experience. This tool exemplifies how expanding our definition of rigor to include different forms of validity—oral tradition, embodied practice, relationship-based knowledge, and place-specific understanding—leads to more effective and legitimate policy outcomes (Althaus, 2020; Papi-Thornton, n.d.). Continual policy evaluation —a practitioner’s guide Authors: Dennis Petrie, Anna Zhu and Pia Andrews, September 2025 How and when do we know that policy is working the way it is intended to? Historically, evaluation of policy in Australia usually occurs as a once-off exercise, long after implementation. Consequently, the magnitude of benefits-–or in some cases, harms—are not known until it is too late, and certain sections of the population may end up worse off than if the policy had never been implemented at all. With recent advances in data availability and statistical methods, this has changed. Now, it is possible to continually evaluate the impacts of policy change and to subsequently adapt or pivot policy where necessary. In other words, we can regularly monitor the performance of policy reforms over time and assess whether or not they are working as intended. What is continual policy evaluation? In many ways, undertaking continual policy evaluation is like running clinical trials for a new drug or medical device. First, we have a social or economic ill that needs to be treated, and so we start to think about possible treatments–in other words, new policy instruments. Once we have a treatment that we think will work, we begin rolling it out to a small section of the population–a “pilot study”–to see if the proposed policy works to effectively treat our identified problem.

During this time, we monitor the study group for potential side effects that we did not necessarily expect to arise; these are our “unintended consequences”. If such negative side effects do arise, we can tailor and temper the policy until it works in the way that was envisaged, while bending the arc of policy towards public good. Furthermore, if we see that our treatment has worked the way it was intended to, we can begin to roll it out to the wider population so that everyone can enjoy the benefits. The most important element of this process is policy adaptation in response to potentially harmful consequences of the policy’s implementation (Swanson et al., 2010). By monitoring policy in real-time, we can pivot the direction or scope of policy when we know that it is not working as intended. To use another metaphor, think of this like a traffic light system. When we observe negative effects of a policy fomenting, we issue a red light: we must stop, wait, and prepare for when we know it will be safe to start moving forward with policy implementation again. When we are not exactly certain about policy impacts, we use an amber light, giving us the time to wait and see how policy impacts evolve before rolling them out to the wider public. Finally, when we see clear and objective signs of policy success we shine the green light to indicate that policy roll-out should move forward, such that all Australians can benefit, noting that what works in one context will not necessarily work in another, meaning that local adaptation will be necessary. Why do we need to do continual policy evaluation? Evaluation is crucial for providing an objective and robust base of evidence to substantiate the decisions made by government. Without a clear cut and evidenced need for policy implementation, policymakers may lose a degree of credibility, or indeed, implement the wrong policies entirely. Regular and consistent policy evaluation is critical to making sure that the policy intent remains valid, that policies work as intended, and that they have not generated any unintended consequences for the people theyhave not generated any unintended consequences for the people they affect.

Policymakers, academics, and the wider public all want to implement the best policy which provides the greatest benefit for the broadest possible cross-section of the population. Similarly, everyone obviously wants to avoid policy which carries negative consequences. However, such negative consequences may not be foreseen when policy is crafted and implemented. Indeed, recent research in both international and Australian contexts has illustrated that often there arise unintended consequences as a result of behavioural, environmental, or other factors that were not necessarily accounted for in the evaluation process (Swanson et al., 2010; Broadway and Zhu, 2023). Continual policy evaluation can fix this problem in several ways. Firstly, by tracking policy impacts in real-time, we can know in the present— not some far-off future—whether or not certain policies actually work. This allows governments (and public institutions) to be more adaptable and responsive to the needs of its people, particularly if a certain policy produces unintended consequences. Moreover, it creates incentives to measure policy outcomes rather than leaving it as a black box for fear of political backlash. Secondly, the “pilot study” phase of evaluation can ensure that only properly tried and tested policies are rolled out to the general public. This is particularly important in a world of over-spending and under-delivering: by conducting smaller-scale studies of the impacts of a policy first, we can ensure that we do not roll out a bad (and potentially expensive) program to the wider population without really knowing its broader effects. Thirdly, continual policy evaluation can give policymakers the confidence to push forward with new ideas even if they may carry some risk because they know that regular monitoring of policy outcomes will provide a “safety net” to catch any unintended consequences. This can revolutionise the policy process, allowing greater scope for novel and innovative approaches to improving the lives of all Australians and people in Australia. So important is the role of evaluation in government that many jurisdictions have now set up registries of policy evaluation to publicly display the evidence base that the choice of policy depends on. The UK, for example, has launched an Evaluation Taskforce and published a registry of policy evaluation documents for public programs, with the ultimate objective of ensuring that “evidence and evaluation sits at the heart of spending decisions”. In the Australian context, there is also clear appreciation of the need for robust policy evaluation practices. In 2021, the government launched the Commonwealth Evaluation Policy, which aims to “embed a culture of evaluation and learning from experience to underpin evidence-based policy and delivery”. This ideal is now further bolstered by the establishment of the Australian Centre for Evaluation (ACE), a central body embedded in the Treasury which collaborates with individual evaluation units across various government departments and agencies. The ACE exists to increase the volume, quality, and use of evaluation evidence in the Australian Public Service. ACE focuses on using causal evaluation tools such as Randomised Controlled Trials (RCTs) or other causal inference methods where an RCT is not feasible. As an example of recent work, the ACE and the Department of Employment and Workplace Relations (DEWR) are testing whether mutual-obligation requirements affect how quickly a study of 30,000 jobseekers find work. For more information see Copley at al. 2025. Such policy evaluations could also be complemented by a practice of continuous evaluation throughout the entire policy lifecycle. For example, it could continuously recruit and/or it could feed back directly into the program as it gradually builds evidence and to ensure optimal policy implementation. How much evidence is “enough” evidence in order to adapt policy? Policy evaluation often faces the problem of being unable to disentangle the impacts of a certain policy from wider changes in the environment or from changes to human behaviour. Indeed, it is very easy to fall into the trap of taking some observable change in behaviour or outcomes to be attributable solely to policy, and to ignore the many spinning cogs of society that operate in the background. Such changes may be entirely spurious or an artefact of the available data and modelling approaches This problem is particularly pronounced when we are testing for many changes at the same time or looking for a certain “signal” of policy impact across a long period of time. This is called “multiple hypothesis testing”. When we make many inferences from data simultaneously there is an increasing likelihood—attributable entirely to chance—that we observe a statistically significant change when no such change has actually occurred. These are “false positives”. In other words, as we test for more and more outcomes, the possibility for such false positives to occur naturally increases, and so we may make incorrect inferences based on our results. Thus, we need some way of knowing—with greater certainty—if a policy actually has worked in the way we intended it to. For that, we need robust causal evidence of its impact. Such assessment requires more advanced statistical techniques and large quantities of data, but advancements across both of these realms have certainly made routine policy evaluation an easier task in recent years. By employing such methods and data, we can ensure that our policy evaluations are reflective of real-world impacts rather than just statistical chance. A more straightforward way to establish if a policy causally affects or contributes to behaviour is to build in impact evaluation principles before the start of a program. As an example, we could measure pre-treatment outcomes before implementing the policy, or build in randomisation or quasi-experimental identification into policies up-front rather than leaving it up to evaluators to find a design that is “good-enough”. What is possible in Australia? It is only in the relatively recent past that the Australian data ecosystem has developed sufficiently to allow for focused analysis of individual- level outcomes on a regular basis. In macroeconomic analysis, real-time tracking and forecasting is often more feasible, as high-frequency data available across a wide range of macroeconomic variables allows for more regular analyses of trends and identification of oncoming challenges. Indeed, this area of research—known as “Nowcasting” has burgeoned since the Great Recession, and was particularly important during the pandemic (Giannone et al., 2008; Bańbura et al., 2013). A recent Australian example is Hartigan and Rosewall (2024). When dealing with public policy that affects individuals on a more micro-level, however, both retrospective and predictive modelling have historically fallen short due to data limitations. For example, the Census of Population and Housing—which provides much of the demographic data required for a considered analysis of socio-economic dynamics in

Australia—is only conducted every five years. Similarly, the Household, Income and Labour Dynamics in Australia (HILDA) Survey—which collects valuable information about economic and personal well-being, labour market dynamics and family life—is available only on an annual basis. Thus, by the time we are able to evaluate policy with the relevant contemporary data, the effects of the policy (whether positive or negative) have likely already bedded in. However, the potential of continual policy evaluation in Australia has been enhanced in recent years thanks to significant advances in the availability of granular administrative data available both through government and industry channels. Of particular note, in 2015 the Australian Bureau of Statistics established the Multi-Agency Data Integration Project (MADIP), a secure data asset which combines and structures multiple streams of data from across the various arms of government. Re-branded in 2023 as the Person-Level Integrated Data Asset (PLIDA), the dataset now offers comprehensive insight into health, education, government payments, income and taxation, employment, and population demographics over time in Australia, with data available from as early as 2006. As such, it is an incredibly useful tool for assessing the efficacy and efficiency of policy over the past two decades, as well as contemporary policy reforms that are being enacted. As of September 2025, there are more than 250 active research projects which draw upon PLIDA data to inform evaluation of policy. Such projects investigate a broad array of outcomes with various data on income and taxation (using ATO data), healthcare usage and outcomes (through Medicare data), and demographic change (using Census data), amongst a variety of others. The Board of PLIDA and the ABS also update and link new datasets as necessary when there is evidence of a potential public benefit. This includes integrating new datasets available via the private sector, such as through the insurance and private healthcare sectors. Continued input of data from both the public and private sectors will further strengthen the data asset in the future, allowing for even more exciting options for policy evaluation to paint a fuller picture of Australian socio-economic dynamics.

Use this chapter

Turn this section into a working policy session, canvas or reference point.