Moderator: Maryam Daneshpour, PhD, MBA, Biotech & Pharma Market, Researcher
Q. What is the biggest industry-level misunderstanding about microbiome or microbial-omics data today?
Pablo - In my opinion, it’s about the various confusions between public and private data. Specifically in microbiome testing, there is so much to gain from a fast go-to-market strategy: that is the moment you start capturing your own private data. This is still highly relevant right now, and right next to it is having a solid strategy to monetise that data, specifically before it loses its value, because the day will come when it becomes commoditised.
Adriel - I would say that correlation does not imply causation. Many microbiome studies that associate microorganisms with health-related conditions are cross-sectional and based on specific populations. Therefore, it is almost impossible to establish the underlying basis of these associations. The microorganism may be the cause, or it may simply be a consequence of the physiological changes that the pathology produces in the body.
Carmen - The biggest misunderstanding at the industrial level is treating microbiome or multi-omic readings as direct and causal biomarkers, assuming that microbiome/omic signals are adequate biomarkers for their purpose by default. Many ‘signatures’ are correlative, highly sensitive to pre-analytical factors (collection, storage), bioinformatic choices, and population context, with limited reproducibility and transferability between different centres. Without a defined context of use, rigorous standardisation, absolute quantification, validated analytical performance, and evidence linking the signal to clinical benefit, these data are weak, and models may appear predictive but cannot be generalised or translated into viable therapies.

Q. Compared with other omics data, what makes microbial omics harder to interpret reliably in clinical programs?
Pablo - With the microbiome, we often try to simplify something immensely complex. We must consider high dimensionality (including phages) and contamination risks, alongside a lack of bioinformatics standardisation, highlighting the need for gold-standard, use-case-calibrated reference datasets. Add to that strong inter-individual variability driven by diet, behavior, and even psychology.
Adriel - There are three main factors: 1) the human microbiome is highly dynamic and prone to change, unlike the genome, for example; 2) it depends on many factors, such as dietary habits or drug intake, which can interfere with interpretation; and 3) the most widely used technique for microbiome sequencing (i.e., 16S rRNA sequencing) provides limited resolution at the species level.
Carmen - Compared with human omics, microbial omics is harder to interpret because signals are compositional, sparse, and highly context-dependent. Taxonomy and function are imperfectly linked due to strain-level variation, horizontal gene transfer, and pathway redundancy. Technical variability (collection, extraction, reference databases, bioinformatics) further limits reproducibility and cross-cohort generalisation, complicating clinically actionable interpretation.
Q. What creates more risk for pharma: unclear data sources, inconsistent curation, or weak quality controls?
Pablo - Inconsistent data curation is a factor that eventually reveals itself and is the hidden showstopper. Poorly structured data, or with insufficient validation, creates a "garbage in, garbage out" cycle that compromises drug discovery. Therefore, applying a professional quality control standard to datasets is the ultimate key.
Q. Why do microbiome biomarkers so often look promising in early studies but fail in prospective or interventional trials?
Pablo - Small sample sizes, driven by high recruitment costs, lead to "noisy" snapshots that overfit specific groups. Without longitudinal tracking, like time series for nutritional bias, and deep sequencing (including viruses), studies capture fragile correlations instead of causation. These early signals may not be enough in trials due to high human variability and, as well, poor data curation.
Adriel - This is a complex question, and there is rarely a single explanation. In most cases, it reflects a combination of factors already mentioned. The proposed biomarker may be specific to the conditions of the initial study (e.g., diet, age, or cohort characteristics), the target population may differ in subsequent trials, or the biological effect may be more modest than initially assumed, resulting in insufficient statistical power in prospective or interventional settings.
Carmen - Microbiome biomarkers often appear to “work” in early studies because small, heterogeneous, and confounded cohorts (diet, medications, geography, sampling, batch effects) enable overfitting and unstable signatures. Many signals are compositional and associative, not causal. In prospective or interventional trials with stricter protocols and new populations, effects shrink and variability dominates, so models fail without strong standardisation and validation.
Q. Based on your experience, how often are clinical failures already “locked in” at the sample collection or sequencing stage?
Adriel - More often than one might expect. In my view, it is essential to start with a clear biological question. However, many clinical studies include microbiome sequencing “just in case,” without a well-defined prior hypothesis. Technical aspects such as sample preservation and transportation, or the choice of sequencing approach (metataxonomy versus metagenomics), are also important, but for me, the most critical point is having a specific question to answer.
Q. How should the industry define “clinical-grade microbiome data,” and in practice, how does it differ from publicly published microbiome datasets?
Pablo - Perhaps, one of the hardest things for us to grasp is that public academic datasets, regardless of their impact (e.g., citations), were not created to meet the highly specific needs of concrete use cases in industry or clinical practice. Their purpose is to serve as benchmarks for research, not as a real Gold Standard. In clinical practice, we need perfectly calibrated datasets, and this requires constant, ongoing data curation and enrichment oriented to the specific use case in turn.
Adriel - There is no universally accepted definition of “clinical-grade microbiome data,” but I look for three things: (1) structured, relevant clinical metadata; (2) a sequencing approach aligned with the objective; and (3) standardised protocols (sampling, DNA extraction, etc.). Public datasets often lack sufficient clinical metadata, limiting interpretability.
Carmen - “Clinical-grade microbiome data” should mean decision-grade, fit-for-purpose evidence generated under processes locked, version-controlled pipelines, reference materials/controls to benchmark performance, auditable provenance, and predefined context-of-use with documented analytical and clinical validation. In practice, many public datasets are discovery-oriented: heterogeneous cohorts and endpoints, evolving methods/databases, limited QC and audit trails, and weaker evidence packages, so they rarely support regulated clinical decisions.
Q. Where does AI meaningfully reduce risk or cost in microbiome clinical programs today, rather than just adding another abstraction layer?
Pablo - The microbiome sits in a bioinformatics “golden age” with ML, digital twins, and synthetic data expanding modeling beyond wet-lab limits. But automation only helps if quality is protected: rigorous input and output validation is essential. Models must stay biologically grounded, automation should accelerate clinical progress, not add complexity.
Adriel - In the integration of clinical metadata with microbiome and other multi-omics data, with the goal of translating population-level discoveries into individualised clinical actions. AI is also valuable for identifying complex patterns and stratifying patient populations, as well as for supporting cohort selection and the systematic use of pre-existing datasets from the literature.
Carmen - AI lowers risk and cost mainly in operational steps, not biological inference. It flags failed samples (contamination, low biomass, index hopping), detects site or batch drift, standardises reporting, and prioritises features for confirmation, reducing rework and delays. It adds risk in patient stratification without a stable ground truth, where models overfit confounders, break after updates, and are hard to lock and validate prospectively.
Q. From your point of view, what regulatory signals do pharma teams often miss early that later become blockers for microbiome-based tools?
Carmen - Pharma teams often miss early regulatory “signals” around the evidence package, not the wet lab. Common blockers include: an ill-defined intended use/context of use (drives product classification and endpoints), lack of locked, auditable algorithms and change-control for microbiome-derived decision tools, and weak plans to demonstrate clinical validity and clinical utility across sites/populations. These gaps surface late as design, trial, and submission blockers.

Q. What lessons should the industry take from the collapse of 23andMe about data ownership, reuse, and long-term trust?
Pablo - First, data assets are subject to obsolescence as they commoditise rapidly. Second, the management of sensitive data requires strict governance due to high ethical and legal risks, alongside threats regarding genetic surveillance. Third, patient/user trust is paramount; losing it can destroy the long-term strategic viability of any data platform.
Adriel - These cases highlight the need for stronger protection of data donors. This requires regulation, but also education: when individuals purchase a service such as 23andMe, they should clearly understand what they are receiving and what rights over their data they are giving up. There is little value in strictly regulating clinical trials if parallel legal frameworks allow sensitive biological data to be collected and reused through consumer services with far fewer safeguards.
Carmen - The 23andMe collapse underscores that “consent once” is not “consent forever.” Data governance must anticipate ownership change, bankruptcy, and secondary use: explicit opt-in for transfers/reuse, clear withdrawal/deletion pathways, and enforceable limits on downstream partners. Trust also hinges on security-by-design and auditable controls; breaches and post-hoc policy assurances rapidly erode willingness to participate, threatening long-term data assets and clinical partnerships.
Q. To what extent do the existing EU and US regulatory and guidance frameworks address privacy risks associated with omics data collected in clinical and non-clinical studies?
Carmen - Not entirely! The EU GDPR treats genetic and health data as “special category” data, but practical gaps remain around secondary use (currently changing with the European Data Space), cross-border exchange, and residual risk of re-identification even after pseudonymisation. The US is more fragmented: HIPAA de-identification may be insufficient for high-dimensional omics, and many non-clinical/consumer and some research contexts fall outside HIPAA. Policies such as NIH-controlled access help, but do not replace harmonised regulation.
Q. For clinical-stage programs, where do current sequencing technologies hit their limits in terms of resolution, cost, or turnaround time?
Adriel - In my opinion, the critical limitation lies in the trade-off between metataxonomy and metagenomics. Metataxonomy (16S rRNA gene sequencing) is highly cost-effective and relatively easy to interpret, but its taxonomic resolution is insufficient for many clinical applications. In contrast, metagenomics offers higher resolution but is more expensive, more time-consuming, and often difficult to translate into clinically relevant insights. While metagenomics should be the preferred option, in practice, it is often challenging to allocate the required resources.
Q. What unique value can small microbiome data companies realistically offer large pharma that big organisations struggle to build internally?
Pablo - We are seeing a continuous rise in the number of Emerging Biopharma Companies, and they are already starting to dominate the market. They are agile, backed by fast-paced external funding, and focused on outlier or niche use cases. While their greatest strength is their speed of development, they could fail without proper compliance oversight. For sure, the data these cohorts generate represents a highly valuable asset for larger corporations.
Q. As the last question, if a pharma company starts a microbiome-enabled clinical program tomorrow, what is the one thing they must get right in the first six months?
Pablo - Without a doubt, the priority is being fast in establishing a robust Gold Standard of reference scientific data, while keeping in mind that it will need regular updates. Although this is public data, it is crucial to optimise and save resources here as much as possible, without compromising quality. Doing so can save months of work and acts as a quality assurance layer for future development, allowing you to accurately evaluate your innovation.
Adriel - Before starting, the biological questions and clinical endpoints must be clearly defined. Within six months, preliminary microbiome data should yield partial insights and guide adjustments (e.g., adding metagenomic or metatranscriptomic layers). Early clarity determines whether data becomes knowledge and ultimately has a clinical impact.
Carmen - In the first six months, they must lock a fit-for-purpose “decision pipeline” end-to-end: a clear context of use, the clinical question and endpoints, and a controlled, versioned workflow from sampling to bioinformatics to reporting with change control and auditability. If this foundation isn’t fixed early, cohort effects, protocol drift, and evolving analyses will invalidate comparability, forcing costly re-sampling, re-analysis, and trial redesign later.
Thank you, Pablo, Adriel, and Carmen, for sharing your thoughtful perspectives and deep expertise on the evolving role of microbiome and microbial omics data in clinical development. Your insights into data integration, translational hurdles, regulatory considerations, and commercialisation pathways have provided valuable insights on both the opportunities and the limitations shaping this space.
We sincerely appreciate your time, expertise, and meaningful contributions to today’s discussion.
