Artificial intelligence in the IVF Lab: A Critical Guide to AI & Machine Learning
Author: Julianna Lamm, MS, Christine Allen, MS, PhD
Artificial intelligence (AI) is everywhere. It powers the recommendation algorithms on your streaming service, the voice assistant on your phone, and increasingly, the tools used in clinical medicine. Now, AI is making its way into one of the most personal areas of medicine: fertility care. Before going further, a word about the word. The American Medical Association (AMA) has argued that in medicine, "AI" is better understood as augmented intelligence - technology meant to extend and support human judgment, not replace it. The tools described here are built to assist embryologists and clinicians and should be judged by their ability to improve relevant targeted outcomes.
This article focuses on where AI is already closest to the bench: the machine learning systems now being deployed in the IVF Laboratory. AI is reshaping clinical practice in other ways too, from large language models that draft patient communications to tools that streamline documentation and scheduling. However, these topics merit their own discussion; here, the focus is on the lab. As a reproductive biologist and machine learning engineer who develops AI tools for fertility care, I want to help our community understand what these systems are, how they are used in IVF today, and what to watch for as this technology continues to evolve.
What Is Artificial Intelligence?
At its core, artificial intelligence refers to computer systems designed to perform tasks that typically require human intelligence - like recognizing patterns, making predictions, understanding language, or supporting decisions. The promise of AI to reduce subjectivity, improve patient outcomes, and personalize treatment protocols has led to its rapid adoption in clinical settings. Before we can discuss its applications and limitations, we must first define and distinguish among the different types of Artificial Intelligence that are rapidly being deployed in our field.
Machine learning is a subset of AI that is particularly relevant to reproductive medicine. While all machine learning is AI, not all AI is machine learning. Rule-based systems and simple automation fall under the AI umbrella but operate on predefined instructions. Machine learning is different: it learns from large amounts of data, detecting patterns and relationships on its own rather than following explicit rules written by a programmer. In fertility care, this could be a system trained on thousands of embryo images to identify features associated with successful implantation.
Most AI tools in IVF today rely on a type of machine learning called supervised learning, in which algorithms are trained on labeled datasets. For example, a dataset of sperm images with morphology labeled as abnormal vs. normal, so that when new unseen images are fed into the algorithm, it can predict whether they are normal or abnormal without requiring manual labeling or human intervention.
A smaller but growing area involves unsupervised learning, which analyzes data without predefined labels to discover hidden patterns, such as identifying patient subgroups that may respond similarly to a given treatment protocol.
It is important to note that AI exists on a spectrum. On one end, relatively simple algorithms automate routine measurements. On the other hand, deep learning models like convolutional neural networks analyze vast image datasets to make nuanced predictions. Understanding where a particular tool falls on this spectrum and how it should be used appropriately matters when evaluating its claims and limitations.
How Is AI Being Used in IVF Today?
AI applications in reproductive medicine have expanded rapidly. A search of the scientific literature reveals over 5,700 publications on AI in IVF and assisted reproductive technology as of 2024, with the vast majority published since 2019. The number of publications has roughly doubled every two years since 2017, reflecting both the promise and the intense interest in this area. Here are the key domains where AI is currently being applied:
1.) Embryo and Blastocyst Assessment
Embryo selection remains one of the most critical and subjective steps in IVF. Traditionally, embryologists visually assess embryos under a microscope, grading them based on morphological features. While this expertise is invaluable, studies show that embryologists agree on grading only about 60-70% of the time, introducing variability that can affect outcomes.
AI tools aim to bring greater consistency and objectivity to this process. Deep learning models, particularly convolutional neural networks (CNNs), can analyze static embryo images or continuous time-lapse video sequences to assess embryo quality, predict implantation potential, and even estimate the likelihood of a live birth. Time-lapse imaging, which captures thousands of images of developing embryos over several days, provides especially rich data for these models. AI can track subtle morphokinetic events - the precise timing and pattern of cell divisions - that may carry predictive value beyond what the human eye can consistently evaluate.
A notable recent development is the emergence of foundation models - large-scale AI systems trained on millions of images - specifically designed for embryo analysis, some achieving meaningful predictive accuracy for outcomes like ploidy status using only imaging data. In September 2025, the FDA granted its first clearance for a machine-learning-based clinical decision-support tool in IVF (Fairtility's CHLOE Blast), marking a significant regulatory milestone for the field.
2.) Sperm Analysis and Selection
On the male side of fertility care, AI is introducing new capabilities to sperm assessment. While manual analysis and computer-assisted semen analysis (CASA) have long been the standard approaches, newer AI-driven tools use machine learning and deep learning to go further. Deep learning models are being developed that can segment and classify individual sperm morphology - including fine structural details like acrosome integrity and head defects - from unstained, live samples at low magnification. Tools like SiD (Sperm ID) take this further, using AI to analyze the real-time motility of individual sperm during ICSI, scoring each sperm based on velocity, trajectory, and head-movement patterns to help embryologists select the optimal sperm for injection. Early studies have shown that higher SiD scores correlate with successful fertilization and blastocyst formation. These systems are also being applied to more challenging clinical scenarios, such as identifying rare viable sperm during surgical retrieval (TESE/mTESE), and at-home semen testing platforms are beginning to incorporate machine learning for consumer-facing fertility screening. Many of these tools are still in early stages of clinical validation.
3.) Automation and Robotics
Perhaps the most ambitious frontier is the development of fully or semi-automated IVF laboratory systems. Several companies are building robotic platforms that aim to automate many of the approximately 200 manual steps involved in creating an embryo in the lab, from dish preparation and sperm processing to ICSI and vitrification. While still in early clinical trials, these systems hold the potential to improve consistency, reduce human error, and eventually make IVF more accessible and affordable.
4.) Personalized Treatment Protocols, Genomics, and Outcome Prediction
Beyond the laboratory, machine learning models are being applied to clinical decision-making. These tools integrate patient-specific metadata - age, BMI, ovarian reserve markers, prior treatment history, stimulation protocols - to predict IVF outcomes and potentially guide treatment decisions. The goal is to move toward more personalized, data-driven care, where stimulation protocols and treatment plans are tailored to each patient's unique profile rather than relying solely on population averages.
AI and machine learning are also increasingly intersecting with genomics. Several companies now use computational models to analyze whole genome sequencing data from embryos, generating polygenic risk scores that estimate an embryo's likelihood of developing common conditions like heart disease, diabetes, and certain cancers. Others are applying machine learning to the sperm epigenome - heritable chemical modifications influenced by lifestyle and environment - to predict fertility status and guide treatment recommendations. In each case, the core value proposition depends on AI's ability to identify patterns across vast amounts of genomic data.
But the fact that AI can generate a score does not mean that the score is clinically sound or actionable. ASRM's Ethics and Practice Committees have concluded that polygenic embryo screening is not yet ready for routine clinical use. This leads us to our next critical point: If the underlying measurement we are asking AI to predict or optimize is itself uncertain, contested, or poorly defined, then no amount of algorithmic sophistication will produce a reliable answer. AI cannot resolve ambiguity in science; rather, it inherits it.
What to Look Out For: A Critical Lens on AI in IVF
While the potential of AI in reproductive medicine is exciting, it is equally important to approach these tools with informed skepticism. Below, we outline key considerations that clinicians, embryologists, and patients should keep in mind:
1.) Bad Data In = Bad Data Out
Consider embryo grading. As we noted earlier, embryologists agree on morphological assessments only about 60-70% of the time. When AI models are trained on these grades as ground truth labels, they inherit that same subjectivity and inconsistency - the algorithm learns to replicate human disagreement, not to transcend it. The ambiguity in the labels becomes baked into the model itself.
A recent study published in Fertility and Sterility (2026) found that commonly used AI models for embryo selection exhibit substantial instability. When the same type of neural network was trained multiple times with slightly different starting conditions, the resulting models produced inconsistent embryo rankings and significant error rates. This raises important questions about the reproducibility and reliability of AI-based embryo selection as currently implemented. This may derive from there being components of inputs that we don't recognize that may bias the assessment independent of the embryos. A non-IVF example was when AI was trained to distinguish huskies from wolves and did so quite accurately based on pictures. However, it was later found that husky images typically had snow in the background, and AI was actually predicting the presence or absence of snow rather than whether the image depicted a husky or a wolf. How cleanly we present embryos without background can potentially affect error rates.
A similar concern applies to models that incorporate preimplantation genetic testing for aneuploidy (PGT-A) data. PGT-A relies on a small trophectoderm biopsy that may not represent the entire embryo, and the clinical significance of mosaic results remains actively debated. If AI models are trained on PGT-A classifications as definitive labels - euploid, aneuploid, mosaic - without accounting for these known limitations, the resulting predictions carry forward those same uncertainties.
2.) Misleading Evaluation Metrics
When evaluating any AI tool, the first question most people ask is "How accurate is it?" In reproductive medicine, the answer can be deeply misleading. Most AI models in IVF are evaluated using metrics like accuracy, AUC (area under the ROC curve), sensitivity, and specificity. These numbers can look impressive in isolation. Some published embryo grading models report AUCs above 0.90. But a high AUC for embryo grading does not necessarily translate to improved clinical outcomes. A model that accurately replicates how embryologists grade is really just replicating an existing subjective standard, the same one we noted earlier, where embryologists agree only 60-70% of the time. If the labels the model learned from are inconsistent, then high accuracy simply means the algorithm has gotten very good at mimicking that inconsistency with no guarantee it will perform similarly on a different patient population or in a different laboratory environment.
A comprehensive review in Patterns (2025) reinforced this finding, noting that performance ceilings across most published models remain below an AUC of 0.85 with significant variability among centers, suggesting that structured data alone may be insufficient to capture the full biological complexity of IVF outcomes. These are meaningful limitations that can be obscured when individual studies highlight their best-performing metric without context.
Equally important are risks inherent to the models themselves. Overfitting, where a model learns the noise and idiosyncrasies of its training data rather than a generalizable pattern, can produce impressive results on internal datasets that collapse when applied to new patients or clinics. Underfitting, by contrast, means a model is too simple to capture the biology it is trying to predict. Neither problem is visible to the clinician or patient who simply sees a score.
When evaluating any AI tool in this field, it is not enough to ask "how accurate is this algorithm?" The more important question is: is the task this algorithm is performing the right one, and are the labels it learned from trustworthy enough to trust the answer?
3.) Transparency and Algorithmic Bias
Many AI tools operate as "black boxes"; they produce a score or prediction, but the reasoning behind that prediction is not always transparent to the clinician or patient. This lack of explainability can make it difficult to understand why a particular embryo was ranked higher than another, complicating clinical decision-making and informed consent. Furthermore, if the datasets used to train AI models are not diverse, if they underrepresent certain racial, ethnic, or age groups, the resulting algorithms may perform less reliably for those populations, potentially exacerbating existing disparities in fertility care.
Looking Ahead
AI is not a passing trend in reproductive medicine. It is a rapidly maturing technology with real potential to improve how we care for patients struggling with infertility, from more objective embryo assessment to personalized treatment planning and automated laboratory workflows.
The responsible adoption of AI in IVF requires rigorous clinical validation, transparent reporting of both capabilities and limitations, equitable development across diverse patient populations, and thoughtful regulatory oversight. Most importantly, it requires that we ask the right questions before we deploy these tools: not just whether an algorithm can perform a task, but whether it should, and whether the data it relies on is solid enough to warrant clinical trust. This is ultimately what the AMA's framing of AI as augmented intelligence asks of us: not deference to the algorithm, but disciplined human judgment working in partnership with it.
As clinicians and embryologists, we encourage you to evaluate AI tools with the same rigor you would apply to any new technology entering your practice. As patients, we encourage you to ask questions about how AI may be used in your care and what the evidence behind it looks like. And as a field, we must commit to holding AI in reproductive medicine to the same evidentiary standards we expect of any intervention that touches a patient's chance at building a family.