Flowchart illustrating the genetic testing process: DNA sample collection, laboratory analysis, data verification, interpretation, and reporting.

The Hidden Risks of using AI to Analyze “Raw DNA”

Why AI analysis can lead to false confidence and harm


By Liza Taylor, MSc, PgDip

Bioinformatics pipeline illustrating DNA analysis and validation process.

When you download a “raw DNA” text file from a consumer genetics company, it can be easy to assume that you have been given the “source code” of your genome.

In reality, that file is almost never raw in the scientific sense. In most cases, it is post-processed output from a bioinformatics pipeline, and it is usually missing the underlying evidence needed to verify accuracy.

That matters because raw data is often used to justify medical decisions and make lifestyle changes, even though it’s often incomplete or unvalidated and sometimes completely wrong.

This article explains what raw data and bioinformatics is and why many consumer raw data files are easy to overvalue.  We will also unpack what proof you should demand from any provider claiming, “bioinformatics included,” and why common AI LLMs cannot do bioinformatics for you.

Bioinformatics is the real product

When you’re given your raw data from a consumer genetics company it is basically noise that needs to be turned into a signal. This is the goal of bioinformatics. It is the set of computational and statistical methods used to turn biological measurements into usable results.

Depending on the test type, bioinformatics involves:

  • Converting instrument output into genotype or variant calls
  • Running quality control checks for identity, contamination, call rate, and coverage
  • Applying thresholds and filters to include, exclude, or flag results
  • Aligning data to a reference genome and coordinate system
  • Annotating variants using scientific databases
  • Applying optional inference steps such as phasing and imputation
  • Deciding how results are reported, including what is shown and how uncertainty is handled

That is why bioinformatics is not a background detail. It is the product. Change the pipeline and you change the answers.

What does “Raw data” mean in genetics?

Raw vs. called data comparison in genetics, microarray scanner intensity measurements and genotype calls.

“Raw data” is not a vibe. In genetics, raw data generally refers to what comes off the instrument before genotype calls or variant calls are finalized. 

Research data standards make this distinction explicit. For example, NIH dbGaP guidance states that for genotyping arrays, raw genotype data like Illumina .idat and Affymetrix .cel should be submitted if available, alongside downstream files that are derived from them. 

If your data is from a SNP microarray (common ancestry style testing)

A microarray scanner produces intensity measurements. Genotypes are inferred later by genotype-calling software.

Illumina’s workflow documentation explains that raw intensity files (.idat) can be converted to genotype call files (.gtc), and that downstream analysis typically works from those converted call files. 

Illumina’s technical documentation shows how cluster files, call thresholds, and locus metrics affect call rate and accuracy in real-world genotyping. 

Illumina’s DRAGEN Array documentation similarly describes .idat as raw intensity input that is used to generate genotype calls. 

If your data is from whole genome or whole exome

Raw data typically refers to FASTQ reads and quality scores, with BAM/CRAM alignments and VCF variant calls as downstream, derived post-processing outputs.

The practical meaning is this: a simple genotype text file is almost never raw in the instrument sense, even if a company labels it that way. 

Why consumer “raw DNA” downloads are easy to overvalue

Most consumer DNA companies offer a downloadable genotype file. A well-known example is 23andMe, which provides documentation about its raw genotype data feature and limitations

Here is the core issue: the typical consumer download is already a simplified genotype report. It is not the raw measurement signal, and it usually does not include the proof needed to verify accuracy.

This is not inherently “bad” for what it was built to do. The problem is when people treat it as decision-grade input for medical, clinical, or lifestyle actions.

A separate issue is completeness. Many consumer tests measure a defined set of markers. That is how microarrays work and how consumer products are scoped. The danger is not that the marker set is defined. The danger is the assumption that the file is comprehensive.

What Does “Raw Data” Actually Mean?  Using 23andMe as a Concrete Example

Consumer DNA test kit and raw data file with validation checkmark.

23andMe is relatively clear about two important points:

  1. Their genotyping detects SNPs (and some other variation types) at a predetermined set of locations called markers
  2. Their raw genotype data has undergone general quality review, but only a subset of markers have been individually validated for accuracy: 

This highlights why a consumer raw download should not be treated as clinical or wellness grade by default. You can have a file full of letters that looks authoritative and still have no way to confirm which parts are solid, which parts are uncertain, and which parts are simply not designed for the use you have in mind.

If you see missing SNPs, it often means the test never measured them. If you see a surprising SNP, it may still require confirmation. The standard accuracy threshold for SNP

genotyping is 99.8%, which means that it’s common for 1 in 500 to be incorrect, especially for rarer variants.

This can be useful for general propensity-based lifestyle testing, but falls short of diagnostic testing for cancers, or carrier testing for BRCA

What Do “Not Determined” Results and Missing SNPs Really Mean? 

DNA sequence grid with magnifier and shield icons illustrating data validation and quality control.

Raw genotype files often include missing calls or “not determined” markers. This is normal in the sense that it is expected in real datasets. It also matters, because it means a call was not confidently made.

23andMe explains raw data browsing and downloading and describes this concept in its raw data navigation guidance.    The deeper point is that a genotype letter pair is not a direct printout from DNA. It is the result of measurement plus an algorithm plus thresholds. Without the underlying evidence and QC context, it is hard to know when a “result” is strong, weak, or misleading.

This becomes especially risky when third-party tools treat missing data as benign, or ignore it entirely, and then make confident lifestyle recommendations.

Verified genetic data analysis and bioinformatics pipelines with quality control checks.

AI chatbots cannot do bioinformatics, and lifestyle genetics is where they fail the most

It is important to separate two goals:

Goal one: Explaining genetics in plain English
Goal two: Performing bioinformatics and calculating defensible genetic risk

Common AI language models can help with the first goal. They can explain terms, summarize concepts, and help you draft questions for a clinician or lab.

They are not designed to be used for goal two.

Chatbot analyzing genetic data with warning sign and statistical charts, illustrating risks of AI bioinformatics.

Bioinformatics requires executing actual pipelines on real files, applying QC, applying statistical methods correctly, and validating outputs. A chatbot generates text. It does not replace compute workflows, evidence, or validation.

When designing our curated poligenic gauges and reports, we meticulously review the scientific literature on each specific trait and accept only the evidence with a high effect weighting and confidence score. This is essential, as AI will  connect a SNP with a specific trait, without looking at the entire genome which might show a contradictory or compensatory gene that ameliorates the effect. 

OpenAI has admitted these issues and has published research on why language models can produce confident falsehoods when uncertain, and why hallucinations remain a persistent reliability challenge.

The narrow exception where chatbots can appear helpful

There are limited scenarios where a chatbot can look “good enough,” especially when the problem is monogenic and tightly scoped, such as counseling-style questions about a small set of well-known variants and established clinical concepts. For example, a study evaluated ChatGPT on common questions related to hereditary gynecologic cancer counseling and found strong performance on that curated question set. 

That does not translate to lifestyle genetics.

Why lifestyle and propensity genetics are different

Lifestyle and propensity genetics often involve many variants, weak effects, population context, scientific uncertainty, and data-science validation. Polygenic approaches, when done properly, are heavy on QC, statistical choices, ancestry calibration, and validation. A well-known guide to PRS workflows makes this complexity explicit: 

In our internal testing at Genemetrics, common AI chatbots show very high error rates on lifestyle-style “genetic risk analysis” prompts:

  • About 90% inaccuracy or hallucination, meaning at least one material error, unsupported claim, or invalid inference in the output
  • Even with tighter constraints, performance has not exceeded about 50% accuracy for complex multi-variant interpretation tasks
  • They tend to break down quickly when asked to reason correctly over more than about 3 to 5 SNPs at a time, especially when evidence quality and study validity must be judged, not just repeated.

Legal and liability: inaccurate genetics has real consequences

A big reason quality and transparency vary across the industry is that the stakes and liability differ by category.

DNA sequence analysis with error and validation symbols, highlighting risks of inaccurate AI genetic interpretation.

When results influence medical and family decisions, legal consequences are real, and enforcement actions and lawsuits follow. For example:

  • The FTC and California obtained an order against CRI Genetics over allegations including misrepresentations about the accuracy of its DNA reports.
  • A class action settlement site describes litigation over allegations relating to Natera prenatal screening test sales, with Natera denying wrongdoing per settlement materials.
  • Time reported on IVF patients suing major labs and testing companies over allegations that PGT-A embryo testing was marketed as highly accurate while being less reliable than implied. 
  • A Connecticut case involved damages tied to an alleged DNA paternity test error, reported byCTPost

The pattern is consistent: when results are marketed as low-stakes, consumers can be left holding uncertainty without realizing it. 

Raw DNA file with question marks and warning sign, highlighting uncertainty and potential inaccuracies.

A consumer “raw DNA” file can be fun. It can support curiosity, genealogy exploration, and learning basic genetics concepts.

What it should not be used for is medical decisions, clinical decisions, or meaningful lifestyle changes, unless your goal is simply to have a small chance of being better than guessing and you are comfortable with the risk of being wrong.

The hard truth is that a genotype text file without a transparent pipeline and without QC proof is not evidence you can rely on. If a provider cannot show you the proof, treat the output as unverified and potentially inaccurate, no matter how confidently it is packaged.

Replace Uncertainty With a Transparent Bioinformatics Pipeline

Your clients trust you to deliver more than automated reports. A consultation gives us the chance to understand your use-case and recommend a white-label structure that turns complex genomic data into clear, defensible insights. Book a confidential call now to see how GeneMetrics can support a more credible DNA offering.

If you’re building or offering genetic insights and want results you can stand behind, talk to us about white-label bioinformatics that is transparent, defensible, and built for real-world use. Educational only. Not medical advice. If a genetic result could change medical care, confirm it through a clinician and a clinical-grade laboratory.

References

NIH dbGaP. Molecular Data Submission Guidelines.
https://www.ncbi.nlm.nih.gov/gap/docs/moleculardatasection

https://www.thermofisher.com/us/en/home/life-science/microarray-analysis/microarray-data-analysis.html

https://help.dragenarray.illumina.com/product-guides/input-files

https://int.customercare.23andme.com/hc/en-us/articles/215983947-Raw-Genotype-Data-Technical-Details

https://customercare.23andme.com/hc/en-us/articles/212883677-How-23andMe-Reports-Genotypes

https://int.customercare.23andme.com/hc/en-us/articles/215983947-Raw-Genotype-Data-Technical-Details

https://customercare.23andme.com/hc/en-us/articles/115004310067-Navigating-Your-Raw-Data

https://openai.com/index/why-language-models-hallucinate

https://www.nateraniptsettlement.com

https://pubmed.ncbi.nlm.nih.gov/38676973

https://www.nature.com/articles/s41596-020-0353-1

https://www.ftc.gov/news-events/news/press-releases/2023/11/ftc-california-obtain-order-against-dna-testing-firm-over-charges-it-made-myriad-misrepresentations

https://time.com/7264271/ivf-pgta-test-lawsuit

https://www.ctpost.com/news/article/jury-awards-2-5-million-shelton-man-dna-18177893.php

Contact us now for a free consultation

Discover more from GeneMetrics

Subscribe now to keep reading and get access to the full archive.

Continue reading