Welcome to SYNTHIA Insight – a series of focused content pieces that bring the science behind SYNTHIA to life. Through interviews with our project partners, we explore views, visions and expertise in synthetic data. Each edition offers an accessible window into the objectives of SYNTHIA and the progress of our work – helping to engage the wider community, spark dialogue, and promote understanding of SYNTHIA’s mission and impact. We invite you to connect with the minds shaping the future of synthetic data.


In SYNTHIA Insight nr. 10, we reintroduce Stella (Styliani-Christina) Fragkouli, Research Associate at the Centre for Research and Technology Hellas (CERTH). In this edition, Stella discusses the motivation and findings behind the publication published in the Frontiers in Bioinformatics, titled: Synth4bench: Generating Synthetic Data for Benchmarking Tumor-Only Somatic Variant Calling Algorithms.


The study explores an important application of synthetic data in cancer genomics: creating reliable ground truth datasets that allow researchers to evaluate the computational tools used to identify genetic mutations.

Cancer develops through the accumulation of genetic alterations, including somatic mutations that arise in cells during a person's lifetime. Accurately identifying these variants is therefore an important part of understanding cancer genomics. However, different computational algorithms can analyse the same sequencing data and produce different results. This creates a fundamental challenge: if the true variants within a real-world dataset are not fully known, how can researchers determine which algorithm is giving the most accurate answer?

Synth4bench addresses this challenge by using synthetic genomics data as a known ground truth. Because researchers control what has been generated, they know which variants are present before asking an algorithm to detect them. This makes it possible to systematically compare algorithms, understand why their results differ, and explore which tools may be better suited to different scenarios.


Watch the video interview:

 

When Different Algorithms Give Different Answers

A central challenge explored in the study is the disagreement between tools used for somatic variant calling. As Fragkouli explains, different algorithms have different mathematical foundations and make different assumptions about the data. As a result, even when they receive exactly the same dataset, their outputs can differ.

“Algorithms that are given the same data produce different results.” This becomes particularly problematic when working with real-world data. If three algorithms produce three different sets of variants with relatively little overlap, researchers face a difficult question: which one is closest to the truth? “At the end of the day, the question is: which one is telling the truth?” Fragkouli explains. “And you cannot tell that because you don't know what's inside your data.” The scarcity of high-quality datasets with known ground truth has long made the robust benchmarking of variant calling tools difficult. Synth4bench was developed to help overcome this limitation by bringing synthetic data generation, variant calling and benchmarking together within one pipeline.


Creating a Ground Truth with Synthetic Data

Synthetic data changes the benchmarking process because the researchers generating the dataset already know what it contains. “Synthetic data gives us the opportunity to have well-characterized datasets. We know what is there before asking the algorithm to make the report, so then we can actually benchmark the algorithm against the ground truth,” Fragkouli explains. Synth4bench generates controlled next-generation sequencing datasets and compares the results produced by different variant callers against this predefined ground truth. The researchers can also systematically change characteristics of the sequencing data, such as sequencing depth and read length, while keeping other parameters unchanged. This allows them to investigate how individual characteristics of a dataset influence an algorithm's behaviour.
This controlled environment provides something that is extremely difficult to achieve with complex real-world datasets: the ability to isolate individual factors and observe how algorithms respond to them.


There Is No One-Size-Fits-All Algorithm

Using Synth4bench, the researchers evaluated five widely used variant callers – Mutect2, FreeBayes, VarDict, VarScan2 and LoFreq. The analysis revealed considerable differences in their performance, with results influenced by the type of variant being detected as well as sequencing characteristics such as read length and depth. For Fragkouli, one of the most interesting findings was not simply that the algorithms disagreed, but what that disagreement tells us about the way these tools model complex biological processes. Different algorithms have different “mathematical hearts”, as she describes them. Some rely on Bayesian inference, others on different statistical approaches, and each weighs components of the data differently. This means that algorithms developed with different assumptions may perform better or worse depending on the research question and type of data being analysed.

“Different algorithms are better for different scenarios,” Fragkouli explains. “It is very important for researchers to know that if you want this type of variant, then this tool works better – or if you want this scenario, go for that tool.” Rather than identifying a universal winner, the results therefore underline the importance of choosing and optimizing tools according to the specific research context. The paper concludes that there is no single solution that performs best across every scenario, and that both sequencing optimization and caller selection are important for maximizing sensitivity and reliability.


Why Some Genetic Variants Remain Harder to Detect

The publication also highlighted the particular difficulty of detecting insertions and deletions, or indels – variants in which small numbers of DNA bases are inserted into or removed from the genome. Across the benchmarking experiments, indels remained more challenging to identify reliably than single nucleotide variants (SNVs).

Fragkouli explains that this is partly due to their more complex nature. Unlike a single nucleotide change, an insertion or deletion can occur in different forms and involve different numbers of bases, creating greater variability for algorithms to interpret. The findings are consistent with previous benchmarking research, but Synth4bench provides a controlled environment in which researchers can examine these difficulties more systematically. In doing so, the study also points towards opportunities for further research into how computational models represent the complex mutational processes underlying cancer.


Beyond Privacy: Another Role for Synthetic Data

Synthetic data is frequently discussed in healthcare because of its potential to support data access and use while protecting patient privacy. Synth4bench demonstrates another important application: using synthetic data as a scientific benchmarking resource. “Synthetic data is a great tool when we talk about algorithm development as well as modelling processes,” Fragkouli explains. Because synthetic datasets can be generated with known characteristics and in large volumes, they can help researchers develop, test and compare algorithms where suitable real-world data may be limited. They can also be used to investigate how computational systems respond to specific data characteristics before those systems are ultimately evaluated against real-world data. This illustrates how synthetic data can support research at different stages, from algorithm development and data augmentation to benchmarking and validation, while helping researchers better understand the strengths and limitations of the tools they use.


Connecting to SYNTHIA's Mission

Within SYNTHIA, this work contributes to the broader goal of developing and validating synthetic data approaches that can support trustworthy healthcare research. While the researchers behind Synth4bench are primarily involved in SYNTHIA's hematological use cases, the benchmarking approach itself has wider relevance across the project, where genomic data can play an important role across different disease areas.

“Our work is very relevant to everyone that is involved in the precision medicine pipeline because we look into the algorithms involved there,” Fragkouli explains. “At the end of the day, our work is to shed light into trusting algorithms within healthcare.” Ultimately, Synth4bench expands the story of what synthetic data can enable. Its value is not limited to generating alternatives to sensitive real-world datasets. By providing controlled ground truth, synthetic data can also help researchers ask whether the computational tools analysing healthcare data are behaving as expected, and understand where and why they may fail.

For Fragkouli, achieving that goal ultimately depends on collaboration. Algorithm developers, bioinformaticians, biologists and medical researchers bring different expertise to the same healthcare challenges. “We have to collaborate with each other and combine our different backgrounds, and we have to go hand in hand in order to produce good results that will help patients, doctors and everyone involved in healthcare.” By bringing these perspectives together, work such as Synth4bench can contribute to more transparent, reproducible and trustworthy computational methods, supporting SYNTHIA's wider ambition to advance the responsible use of synthetic data in precision medicine.


Listen to the podcast:


For the full publication: