Software testing has changed considerably over the past five years. Automation has become an established part of software delivery, AI can now generate test cases in seconds, and quality engineering teams have more tools available than ever before. At the same time, organisations are releasing software more frequently, applications are becoming more complex, and testing teams are expected to keep up without necessarily receiving more resources.

But has all this progress made software releases safer? Or have we simply become better at producing and executing more tests?

To understand what has changed, we examined several editions of three major industry research series: the World Quality Report 2025–2026, published by Capgemini, Sogeti, and partners; Tricentis' Quality Transformation Report; and PractiTest's State of Testing. We looked at their findings across the 2021–2026 period, paying particular attention to recurring questions, changes in reported adoption, and differences between industries and respondent groups.

The reports approach software quality from different perspectives, which makes comparing them both useful and complicated. One focuses on enterprise technology strategy, another on the experiences of testing professionals, and the third on delivery pressure and the consequences of poor software quality. They do not always measure the same things, and a difference between two percentages does not automatically mean that the research disagrees.

Nevertheless, several patterns emerge. AI adoption has accelerated much faster than enterprise-wide implementation, automation growth has not eliminated testing workloads, and organisations continue to struggle with the relationship between testing activity and business outcomes.

What five years of research can (and cannot) tell us

Before comparing the findings, it is worth understanding who participated in the surveys. This is particularly important when looking at AI adoption, because an organisation investing in a technology and a tester using that technology every day are two different things.

The World Quality Report primarily surveys senior IT and quality engineering decision-makers, often from large enterprises. Its 2024 edition included 1,775 IT executives, making it particularly useful for understanding organisational investments, technology priorities, and quality engineering transformation. These respondents can describe what their organisations are planning or implementing, although that does not necessarily mean every initiative has reached the teams responsible for executing tests.

PractiTest takes a different approach, surveying QA engineers, automation specialists, testing managers, and other testing professionals. Its findings provide a closer look at everyday testing practices, professional responsibilities, and the challenges teams encounter. Tricentis combines perspectives from technology executives, engineering leaders, developers, and QA professionals. Its 2025 report surveyed approximately 2,750 respondents across 10 countries and five industry verticals, with particular attention to release pressure, software failures, and their business consequences.

There is also a historical difference between the publishers. The World Quality Report and State of Testing have been published for many years, making it possible to examine how certain topics have developed across editions. Tricentis' Quality Transformation Report is a newer series, so its 2025 and 2026 findings offer a shorter comparison rather than a continuous five-year history.

This distinction matters throughout the analysis. A percentage describing organisations experimenting with AI cannot be compared directly with one describing testers who use AI for a specific task. Likewise, a change between annual editions may partly reflect differences in the survey population or question wording rather than a change in the industry itself.

The strongest conclusions therefore come from examining comparable findings within each series first, then looking at where the different reports point towards similar problems.

1. AI adoption accelerated, but implementation has been less straightforward

The most visible change in software testing research has been the growing importance of artificial intelligence. Earlier editions of industry reports concentrated heavily on traditional automation, Agile and DevOps adoption, testing environments, and the maintenance of increasingly complex systems. Those subjects remain relevant, but generative AI has become one of the dominant themes in recent research.

The World Quality Report illustrates how quickly this happened. Across its 2023–2024, 2024–2025, and 2025–2026 editions, the reported share of organisations not exploring or adopting generative AI in quality engineering moved from 31% to 4%, before increasing to 11%.

EditionOrganisations not exploring or adopting GenAI
2023–202431%
2024–20254%
2025–202611%

The initial movement was considerable, with a 27-percentage-point decline between the first two editions. The subsequent increase is more difficult to interpret. It could reflect a more selective approach to AI adoption, but changes in survey composition or definitions may also have influenced the result. It would therefore be misleading to describe it as evidence that organisations began abandoning AI.

The 2025–2026 report provides a clearer distinction between experimentation and implementation. It found that 89% of organisations were piloting or deploying generative AI in quality engineering, while only 15% had reached enterprise-scale deployment. Data privacy concerns, integration complexity, and skills gaps were among the most frequently reported obstacles, cited by 67%, 64%, and 50% of respondents respectively.

In other words, experimenting with AI has become relatively common, but integrating it across an organisation's quality engineering processes remains considerably more difficult. This is especially relevant in enterprise environments, where testing often involves sensitive information, legacy applications, multiple integrations, and established governance requirements.

Test creation is getting more attention than risk identification

PractiTest's 2026 findings provide another perspective. The report found that 76.8% of respondents reported AI usage within their organisations. Although this is lower than the World Quality Report's 89%, the figures measure different populations and forms of adoption, so they should not be treated as contradictory.

The more interesting comparison is between the activities for which AI is being used.

AI applicationReported usage
Test-case creation69.6%
Script and test maintenance59.6%
Risk identification19.9%

AI is being used much more frequently to create and maintain testing assets than to identify which risks deserve attention. That does not mean the generated tests are unnecessary, but it raises an important question about where organisations expect AI to create value.

A team that previously spent several days preparing 200 test cases may now be able to generate an initial set in minutes. This can reduce repetitive work, particularly when requirements change frequently. However, generating tests faster does not establish that the most consequential failure scenarios have been identified. If a specification overlooks an important business rule or contains ambiguous requirements, generating more tests from that specification may reproduce the same limitations.

The distinction is between making an existing process faster and improving the decisions behind that process. As we discuss in our article on intelligent test automation, AI can support test creation, execution, and maintenance, but those capabilities still need to be connected to a testing strategy.

The research suggests that AI adoption has progressed further in activities that increase testing output than in those that help determine testing priorities.

2. Automation keeps expanding, but testing teams are not necessarily becoming less busy

Automation was already a central part of software testing before the recent wave of generative AI. Over the past five years, the conversation has increasingly moved from whether organisations should automate testing to how they can maintain larger suites, support more frequent releases, and reduce the effort required to keep automated tests reliable.

There is an important distinction between automating individual activities and reducing the overall workload of a testing team. Automated execution may eliminate repetitive manual checks, but scripts still require maintenance, test environments need management, and failures need investigation. Meanwhile, development teams continue introducing new functionality, integrations, and dependencies (all of which create additional scenarios that may need testing).

PractiTest's 2026 research shows how differently automation is progressing across industries.

IndustryRespondents reporting increased automation
Finance and insurance67.1%
Internet and technology54.4%
Transportation51.9%
Health services50.0%
Retail35.7%

These percentages represent respondents reporting increased automation, not the proportion of tests already automated within each sector. Financial services reporting 67.1%, for example, does not mean that financial institutions have automated approximately two-thirds of their tests.

Even with that distinction, the differences between industries are substantial. Finance and insurance respondents were almost twice as likely as retail respondents to report automation growth, although the survey does not establish exactly why. Regulatory requirements, organisational resources, technology infrastructure, and commercial priorities may all contribute.

What makes the results more interesting is that automation growth appears alongside increasing workloads. In finance and insurance, 68.3% of respondents reported growing workloads. In healthcare, the equivalent figure was 65.4%.

These findings do not prove that automation creates additional work. However, they challenge the assumption that more automation necessarily means testing teams have less to do. An organisation can automate more checks while simultaneously expanding the number of applications, environments, and releases its QA team is responsible for.

The same issue appears when teams evaluate automation primarily through the number of automated tests. A regression suite can grow substantially without becoming easier to maintain or more effective at detecting serious failures. Our guide to regression testing examines why selecting relevant tests and maintaining reliable execution are just as important as increasing the size of the suite.

For organisations investing in automation, the more useful questions concern maintenance effort, execution reliability, feedback speed, and whether automation improves the team's ability to validate important workflows before release.

3. Faster releases have not eliminated incomplete testing

Tricentis' research brings the discussion closer to software delivery itself. Rather than focusing primarily on technology adoption, it examines what happens when organisations try to release software quickly while maintaining acceptable quality standards.

The 2025 Quality Transformation Report found that 63% of respondents reported shipping code changes without fully testing them. Pressure to accelerate release cycles was cited by 46%, while 40% reported accidental deployment of untested code. The research also found that 45% prioritised improving delivery speed, compared with 13% prioritising software quality.

These findings describe a problem that many development teams will recognise. Businesses want new features delivered quickly, customers expect continuous improvements, and engineering teams are expected to maintain increasingly complex applications without slowing down releases.

The 2026 Tricentis report suggests that the issue has persisted, with approximately 60% of respondents continuing to report deployment of untested code. Compared with 63% in 2025, the difference is relatively small, although the underlying question wording and respondent composition would need to be checked before treating it as a precise year-on-year improvement.

An important limitation is that these figures do not tell us how frequently individual organisations release incompletely tested changes, or how much risk those releases introduce. An organisation that occasionally excludes low-priority regression checks and one that regularly deploys critical functionality without adequate validation could both appear within the same statistic.

This is why the question of release confidence needs more context than a simple distinction between tested and untested code.

Imagine a team preparing to release a change to an online banking application. The regression suite contains 30,000 automated tests, and all of them pass. The execution dashboard is green, no failures require investigation, and the team has met its established testing targets.

From an execution perspective, everything looks good. But if the suite does not adequately cover a critical payment authorisation scenario, the passing results may provide more confidence than the evidence actually supports.

The tests have passed, but that does not necessarily mean the most important risks have been addressed.

This distinction also matters when evaluating software delivery performance. In our article on DORA metrics in software testing, we explain how testing can influence deployment frequency, lead time, production failures, and recovery. These outcomes provide a broader perspective than simply counting executed tests or measuring automation percentages.

4. We measure testing activity more easily than its business impact

One of the recurring questions raised by the reports concerns what organisations choose to measure.

PractiTest's 2026 research found that 56% of respondents reported being measured on test coverage, while only 4.5% reported Net Promoter Score as a testing performance metric. These measures are not equivalent, and NPS would not necessarily be appropriate for every testing team. Nevertheless, the comparison illustrates how much easier it is to measure what happens within testing than to connect those activities to customer or business outcomes.

Test coverage is useful because it helps teams understand which requirements, components, or workflows have been validated. Execution reports show which tests passed, which failed, and where further investigation is needed. These are essential parts of managing a testing process, particularly when organisations maintain large automated suites.

The limitation appears when coverage or execution results are treated as sufficient evidence of release safety.

A team may achieve 95% requirements coverage, but that percentage does not tell us whether the remaining 5% contains functionality capable of causing serious financial or operational damage. Similarly, executing thousands of automated tests may demonstrate that a large amount of functionality behaves as expected, without establishing whether the most consequential failure scenarios were included.

Tricentis' 2025 findings make the business implications clearer. According to the report, 42% of respondents believed poor software quality cost their organisations at least $1 million annually. These are respondent estimates rather than independently audited financial losses, but they illustrate why quality cannot be evaluated exclusively through engineering activity.

The consequences of a defect depend heavily on what the affected software does. A visual inconsistency in an internal dashboard and a failure in payment processing may both count as one defect, but their potential impact is very different. The same applies to test cases: two tests may require similar execution effort while protecting workflows with substantially different business consequences.

This is where the distinction between test coverage and risk coverage becomes important. Test coverage describes what has been tested, while risk coverage considers whether the testing effort addresses the failures most likely to cause serious consequences.

Our guide to risk-based testing explores how teams can use business impact, technical complexity, and likelihood of failure to prioritise testing. The objective is not to replace conventional coverage metrics, but to give them more context.

5. Industry and company size change what the numbers mean

One of the limitations of global software testing statistics is that they can hide substantial differences between organisations. A financial institution, a healthcare provider, and a retail company may use similar testing tools while facing very different operational constraints.

PractiTest's 2026 sector breakdown provides useful evidence of those differences, although it represents a snapshot rather than a complete five-year trend for each industry.

Financial services: More automation, but security remains a concern

Financial services reported the strongest automation growth among the sectors examined, with 67.1% of respondents indicating an increase. At the same time, 54.9% identified security as a primary obstacle to AI adoption.

This combination is particularly relevant for banking and insurance organisations, which frequently operate complex applications involving sensitive information, legacy systems, and regulatory requirements. Introducing AI-assisted testing into these environments may require additional controls around data handling, access, traceability, and the reliability of generated outputs.

The sector also illustrates why testing priorities need to reflect business consequences. Updating a notification preference and authorising a high-value transfer may both be part of the same application, but a failure in the latter could create significantly greater financial, operational, and regulatory exposure.

For organisations testing banking applications, these differences are especially important when validating complete transaction journeys. Our guide to SWIFT payment testing examines the requirements involved in validating payment messages, routing, integrations, and confirmations across connected banking systems.

Healthcare: Growing workloads without equivalent budget growth

Healthcare respondents reported a different combination of pressures. In PractiTest's 2026 research, 65.4% experienced increased workloads, while only 11.5% reported budget growth. Another 38.5% identified AI tool complexity as an adoption barrier.

These findings suggest that healthcare testing teams may be managing expanding responsibilities without comparable increases in resources. In environments where applications process sensitive patient information or support critical workflows, reducing testing effort also requires careful consideration of validation requirements and the consequences of failure.

Automation can help reduce repetitive work, but it does not remove the need to review test results, validate requirements, and demonstrate that critical functionality behaves as intended. When budgets are constrained, the challenge becomes deciding where additional testing effort provides the greatest value without compromising mandatory validation activities.

Retail: A different business case for automation

Retail respondents reported lower automation growth, at 35.7%, and the same proportion identified return on investment as an obstacle to AI adoption.

It would be misleading to conclude that retail organisations are less mature in their testing practices based on these figures alone. The survey measures reported growth rather than total automation coverage, and the organisations represented may differ in size, technology infrastructure, and business priorities.

Nevertheless, the findings suggest that automation investment is not driven by identical considerations across industries. For some organisations, regulatory requirements may strongly influence testing decisions. For others, the business case may depend more heavily on measurable efficiency gains, maintenance costs, and the prevention of customer-facing disruptions.

Large enterprises and small teams use AI differently

Organisational size introduces another important difference. PractiTest found that AI adoption reached 81.7% among respondents from organisations with more than 10,000 employees, compared with 70.6% among those from organisations with 1–10 employees.

The applications of AI also varied.

AI applicationLarge organisationsSmall teams
Script maintenance65.5%41.7%
Test planning31.0%50.0%
Risk identification12.1%33.3%

Large organisations were more likely to report using AI for script maintenance, while smaller teams reported greater use of AI for planning and risk identification.

One possible explanation is that larger enterprises have more extensive existing automation infrastructure and therefore spend more effort maintaining it. Smaller teams may have fewer specialised resources and see greater value in using AI to support planning. These interpretations are plausible, but the survey does not establish them as causes.

What the findings do demonstrate is that a single AI adoption percentage can conceal very different practices. Two organisations may both report using AI in testing while applying it to entirely different problems.

6. Where the reports agree, and where comparisons become misleading

Read individually, the three research series can appear to describe different industries. The World Quality Report shows widespread enterprise interest in AI and quality engineering transformation. PractiTest shows testing professionals adopting new technologies while continuing to experience workload pressure. Tricentis highlights incomplete testing and the business consequences associated with software failures.

These findings are not necessarily contradictory. An organisation can invest heavily in automation while its QA team remains under pressure. A company can introduce AI without having integrated it into every testing workflow. And a team can execute a larger regression suite while still experiencing production defects, particularly when important failure scenarios are missing from its coverage.

The reports examine different stages of the same process: investment, implementation, everyday usage, and business outcomes.

There are, however, limits to the comparison. The World Quality Report has a long publication history, while the Quality Transformation Report provides a much shorter series. PractiTest offers valuable historical material, but its respondent population differs from the enterprise-focused research. Question wording, sector composition, and definitions can also change between editions.

This is particularly important when interpreting adoption rates. The World Quality Report's 89% figure and PractiTest's 76.8% figure do not prove that executives are more optimistic about AI than testing professionals. They describe different populations and measurements, so the difference cannot be attributed to respondent seniority without additional evidence.

The strongest common finding is therefore not a single percentage. It is the relationship between growing testing capabilities, continued delivery pressure, and the difficulty of demonstrating how testing reduces business risk.

7. The next challenge is deciding what deserves testing

For years, organisations have invested in making testing faster and less dependent on repetitive manual execution. That remains valuable, particularly for regression testing, where reliable automated checks can provide frequent feedback as applications change.

But as AI makes it easier to generate tests and automation makes it easier to execute them, deciding which tests are necessary becomes increasingly important.

If a team can generate 200 test cases in minutes, should it execute all 200? If an organisation already maintains 30,000 regression tests, should its next objective be to reach 40,000? And if every automated check passes, how does the team determine whether the application is sufficiently safe to release?

These questions cannot be answered through test quantity alone.

Risk-based testing provides a way to connect testing priorities to the potential impact and likelihood of failure. Instead of treating every requirement as equally important, teams assess which workflows could cause the greatest financial, operational, regulatory, or customer consequences if they failed.

For a banking application, that might mean prioritising payment authorisation, account access, and transaction processing. For a healthcare system, it might involve patient data integrity or critical clinical workflows. The objective is not to ignore lower-risk functionality, but to ensure that limited testing resources are allocated with an understanding of what is at stake.

This is also where AI could become more useful than a tool for generating additional scripts. If it can help teams analyse specifications, identify missing scenarios, and evaluate potential consequences, it can contribute to the decisions that shape the testing strategy rather than simply accelerating execution.

At TestResults, this is the idea behind Leonidas, which analyses specifications, identifies test cases, and scores them according to potential damage and frequency. Combined with automated execution, this connects the decision about what deserves testing with the ability to validate those workflows reliably.

The difference is important. Generating more tests may increase coverage, but identifying which tests address the greatest risks can make that coverage more meaningful.

Frequently asked questions

1. How has software testing changed over the past five years?

Software testing has become increasingly automated, with AI playing a much larger role in test creation, maintenance, and execution.

According to the World Quality Report 2025–2026, 89% of surveyed organisations were piloting or deploying generative AI in quality engineering, although only 15% had reached enterprise-scale deployment. The shift has been substantial, but faster testing has also introduced new challenges around maintaining automation, managing complexity, and determining which tests provide the greatest value.

2. Does more test automation lead to better software quality?

Not necessarily. Automation can improve execution speed, consistency, and coverage, but its effectiveness depends on which scenarios are being tested. A large automated regression suite may still overlook critical business workflows, particularly if tests are selected primarily to maximise coverage rather than reduce risk. The important distinction is between increasing the number of automated tests and improving the organisation's ability to detect consequential failures before release.

3. Why do software testing industry reports show different statistics?

The World Quality Report, Tricentis' Quality Transformation Report, and PractiTest's State of Testing survey different populations and measure different aspects of software quality. For example, an enterprise executive may report that their organisation has adopted AI, while an individual tester may not yet use AI in their daily work. Differences in geography, industry, company size, survey wording, and respondent seniority can also influence the results. Percentages should therefore only be compared directly when they measure equivalent behaviours in comparable populations.

4. What should organisations prioritise in their software testing strategy?

Organisations should look beyond test counts and automation rates to understand which business risks their testing actually addresses. This means identifying critical workflows, assessing the potential impact and likelihood of failure, and prioritising testing accordingly. Combining risk-based test selection with reliable automation helps teams make more informed release decisions, particularly when time and resources are limited. The objective is not simply to execute more tests, but to have stronger evidence that the most important risks have been addressed.

We have more tests than ever. But do we have more confidence?

Five years of software testing research show an industry that has become remarkably good at increasing its output. We can automate more workflows, generate test cases in seconds, and execute regression suites that would have required enormous amounts of manual effort not that long ago. Yet the same reports suggest that some of the industry's most persistent problems have not gone away. Teams are still under pressure, organisations still release software without completing all their planned testing, and the cost of poor quality remains significant.

Perhaps the problem is that we have spent so much time improving how we test that we have paid less attention to how we decide what deserves testing in the first place. We measure coverage, execution speed, and the number of automated checks because these are relatively easy things to count (and they look particularly good on a dashboard). What is considerably harder to measure is whether those checks protect the parts of an application where failure would actually hurt the business.

The reports cannot prove that more automation has failed to improve software quality. But they do expose a gap between the capabilities organisations are investing in and the outcomes they need to demonstrate. And with AI making it easier than ever to generate additional tests, that gap deserves more attention. Otherwise, we risk building increasingly impressive regression suites without becoming much better at understanding what they protect.

Maybe the next major improvement in software testing will not come from another tool that helps us execute 10,000 more tests. Maybe it will come from being able to explain why we needed those tests, what would happen if the scenarios they cover failed, and which risks are still worth worrying about.

Because 30,000 passing tests can tell you that everything you checked is working. They cannot tell you whether you checked everything that matters.