Last week we put a price on the software. This week, let's open it up.
Imagine a bitewing entering an AI system. A colored outline appears around a suspicious proximal surface. Somewhere behind that outline, people made decisions about what counts as disease, which patients the model could learn from, how much error to tolerate, and what the dentist should do with the result.
Those decisions determine whether the outline deserves attention.
The same is true of software that traces a mandibular canal, predicts periodontal measurements, or prepares a claim. They may share a login and a subscription. Underneath, they can be very different machines.
A useful way to understand them is to follow the build: define the job, assemble the evidence, train a model, test it on genuinely unfamiliar cases, and watch what happens when people use it. Then we can look at why Retrace is technically interesting, and how much of its broader promise the public evidence supports.
Start with a question precise enough to get wrong
Suppose the job is to identify a radiographic finding consistent with proximal caries. The input might be a bitewing. The output might be a probability and a marked surface.
That specification leaves important clinical questions open. Is the lesion active? Is it cavitated? Has it progressed? What does the examination show? What management is appropriate for this patient?
A team has to decide which question it can actually label and evaluate. Otherwise a model trained to reproduce a reader's radiographic judgment can quietly acquire a much bigger description: software that determines treatment need.
The training target is the answer the developers teach the system to predict. A clinical measurement, an expert interpretation, a future emergency, and a paid claim are different targets. Getting the target right comes before choosing the neural network.

Figure 1. Focused dental AI products and context-rich systems differ in scope and architecture. Complexity does not establish clinical quality, and these examples do not describe any one vendor’s deployed technology.
What happens during training
An image model converts pixels into numerical features. Early layers might respond to edges or textures; later representations combine information into patterns useful for the task. An encoder produces these representations. A task-specific head converts them into an output.
A classification head might say whether a finding is present. A detector also localizes it. A segmentation head assigns a label to each pixel, or each three-dimensional voxel in a CBCT volume. A landmark model estimates positions from which a distance can be calculated.
During training, the model makes predictions on batches of examples. A loss function assigns a penalty for disagreement with the training labels. Backpropagation calculates how changing the model's adjustable numbers, called parameters, would change that penalty. An optimizer makes small adjustments. Repeating this process teaches the network patterns that reduce its training error.
The training objective matters. A penalty dominated by normal surfaces can reward a model that misses uncommon disease. Sampling more positive cases or weighting their errors more heavily can help, but those choices also affect how its scores behave. They need evaluation.
A convolutional neural network builds features with learned local filters. A vision transformer divides an image into patches and uses attention to combine information across them. Both are architectural tools. The name alone says little about the quality of the labels or the consequences of a mistake.
For CBCT, the engineering also includes voxel spacing, crop size and physical coordinates. Reducing resolution makes a scan easier to process, but can remove the detail needed to locate a narrow boundary. A coarse model can first find the relevant region; a finer model can examine it more closely. A published oral-surgery segmentation system used separate networks for bone, neural-tube and tooth extraction, illustrating why one product may contain several specialized models. Its segmentation results do not establish fewer surgical complications. 2024 primary study
The best starting point is a strong comparison
Building from scratch means learning the parameters without a pretrained starting point. Transfer learning starts with an existing model and adapts some or all of it to the new task.
For a team with limited labeled dental data, I would begin by comparing a strong task-specific model with a pretrained vision model. Keep the patients, annotation budget and evaluation rules the same. A more elaborate system should earn its complexity.
The nnU-Net work is instructive here. It showed how much segmentation performance depends on systematically configuring preprocessing, architecture and training for the dataset. Clever network design is only one part of the build. nnU-Net
Pretraining can also use images without a dentist labeling every finding. In self-supervised learning, the training exercise comes from the data itself. A masked autoencoder hides image patches and learns to reconstruct them; the learned representation can then be adapted to a labeled clinical task. Reconstruction ability by itself establishes no diagnostic competence. Masked autoencoders
A foundation model is intended to provide reusable representations across tasks. It need not generate language. The 2025 DentVFM preprint describes dental vision models pretrained on roughly 1.6 million radiographic images. That is an interesting research direction, with benchmark evidence. It is still a preprint, and those results do not establish prospective patient benefit. DentVFM
The question for a builder is what that pretraining adds under a fair comparison. Does it need fewer labels? Handle a new scanner better? Improve difficult boundaries? A large image count becomes useful when those questions have answers.
Build the labels as carefully as the model
A dataset starts with clinical records that were usually collected to deliver care. Turning them into a training set creates another layer of decisions.
For our bitewing model, a labeling manual should define positive, negative, uncertain and ungradable surfaces. Readers should work independently before disagreements are resolved through a specified process. Their disagreement is information worth retaining.
The term reference standard is helpful. It identifies what the model is being compared with without implying perfect truth. Expert consensus, operative findings and longitudinal follow-up answer different questions and have different limitations. The CLAIM reporting guidance asks developers to describe annotators, disagreement handling, data partitions and external testing. CLAIM 2024
An absent mention in a report is a weak basis for declaring disease absent. A procedure code says care was recorded. A paid claim says a payment process reached a particular outcome. Neither automatically supplies the clinical label we wanted.
Count unique patients as well as images. Audit restorations, implants, mixed dentition, artifacts and ordinary negative cases. A collection of beautiful textbook pathology can train a model for a world the practice rarely sees.
Keep the exam answers out of the training room
A model can look excellent because its test is too easy.
For an evaluation of new-patient performance, keep every image, crop and visit from one patient in the same partition. Otherwise training and testing may share the patient's anatomy, restorations or near-duplicate images. Thousands of CBCT slices do not amount to thousands of independent patients.
Use a training set to fit parameters, a validation set to choose settings, and a locked test set for the final assessment. If developers repeatedly choose models by inspecting test results, that test has become another tuning set. Pretraining data need scrutiny too: a test case can be familiar even when no disease label accompanied it.
Then ask three separate questions. Does it work on new patients from familiar sources? At a different practice or scanner? In a later calendar period? Patient-disjoint, external-site and temporal testing address different uncertainties. A single random split cannot answer all three. CLAIM 2024
Longitudinal prediction adds a hard boundary: the prediction time. If we are estimating next year's risk, the input must contain only information available today. Later treatment codes or postoperative notes can leak the answer backward into the model.
Read the errors at the threshold you will use
A model's score becomes a clinical flag when it crosses a chosen threshold. Lowering that threshold usually catches more disease and creates more false positives. The useful operating point depends on the job and the cost of each mistake.
Consider a hypothetical evaluation of 1,000 surfaces, with disease in 5%. At 90% sensitivity and 90% specificity, the model finds about 45 true positives and flags 95 healthy surfaces. Only about 32% of its positive flags are true positives.
Those are illustrative numbers, not a product result. They show why a headline accuracy claim needs prevalence, a threshold and a stated unit. Surface-level performance cannot quietly become patient-level performance.
Calibration asks another question: when a system assigns a probability near 80%, is the event observed about 80% of the time in that population? A model can rank cases well while being overconfident. Calibration can be adjusted on held-out data, but it needs checking again when the population or equipment changes. Calibration methods · Uncertainty under dataset shift
For segmentation, average overlap is insufficient. A canal contour can agree with most of a reference tracing while making a locally important error. Evaluate boundary distances and correction burden where errors matter clinically. Also examine difficult subgroups. Strong aggregate performance can conceal weak performance in a clinically meaningful subset. Hidden stratification
A model that sometimes declines to answer can be useful. Measure which cases it declines and which serious errors remain among the answers it accepts.
Where language models and patient context belong
A large language model predicts text from its context. That can help turn an encounter into a draft note or organize evidence for a claim. A reliable application needs more than the generator.
For a claim narrative, I would start with retrieval of the actual chart entries and applicable policy, extract the relevant facts into structured fields, validate tooth numbers and required attachments, and then generate a reviewable draft. Each factual statement should trace back to a source. Missing evidence should remain missing.
Retrieval-augmented generation supplies selected documents when an answer is produced. Fine-tuning changes model parameters. These solve different problems; retrieval makes sources updateable and inspectable, while fine-tuning can adapt behavior. Neither guarantees factual accuracy. Original RAG research
A multimodal system combines different kinds of input, such as radiographs, charting and notes. Separate encoders can turn each into features that are combined before prediction. Another design combines the outputs of separate models. Shared representations may capture useful relationships; modular designs can make failures easier to isolate.
The extra inputs must justify themselves. An ablation tests the system with a component or input removed. If history improves a risk model, does the improvement persist at a new practice and on later patients? If payer identity improves claim prediction, the model may be learning reimbursement behavior. That is useful for a different purpose.
Retrace gives us a concrete case to examine
Disclosure: Ali Sadat, Retrace’s founder, is a former student of mine.
Retrace is interesting because its public work connects image analysis with documentation and the machinery of claims. Three kinds of evidence need to stay separate: a published experiment, disclosed technical designs, and current product descriptions.
In a 2022 Journal of Dentistry study, Kearney and colleagues, including Retrace's Ali Sadat, tested inpainting before predicting clinical attachment level, or CAL. CAL is the clinical probing distance from the cementoenamel junction to the base of the sulcus or periodontal pocket. The study used a generative adversarial network, or GAN, with partial-convolution inpainting, followed by prediction models adapted from DeepLabV3+ and DETR. These draw on segmentation and transformer-based detection architectures, respectively. Training included 80,326 images from 9,264 patients. Reported mean absolute error was 1.04 mm with inpainting versus 1.50 mm without it. 2022 study
The retrospective study used selected institutional records with a six-month matching window and excluded anterior teeth, third molars, inconsistent records and high CAL measurements. The authors report withholding 10,687 images and 40,077 CAL measurements from data scientists for once-only, blinded testing; final predictive analysis included 1,911 unique patients. Explicit patient-disjointness across every split and external-site testing could not be established from the methods reviewed. Study record
A GAN trains a generator alongside a discriminator that learns to distinguish generated from real examples. Their competing objectives encourage plausible generated content. Inpainting uses observed context to estimate missing regions; partial convolutions account for which parts of the input are available.
The engineering result is worth attention: preprocessing improved the downstream estimate in this experiment. The clinical distinction is equally important. Generated pixels are estimates. They are not measurements acquired from the patient, and a plausible completion cannot establish unseen anatomy.
Before using such a method in care, I would want the original image preserved, synthetic regions clearly identified wherever displayed, and validation of the full downstream decision. I would also compare inpainting with a model given the incomplete image and an explicit missing-region mask. Does generating anatomy add useful information beyond teaching the predictor how to handle missing data?
Those are proposed tests. The paper does not establish that synthetic images are currently shown to clinicians in a commercial product, nor does its average error establish improved treatment decisions.
The longitudinal idea deserves a separate test
Retrace's published patent application US20220012815A1 describes a broader architecture. One stage extracts findings from images and documentation. Later models use those findings and other information to estimate documentation deficiencies, claim adjudication and payment-related outcomes. Figure 29 combines image-derived information with clinical metadata and payer identity. Figure 30 describes appointment representations processed in temporal order through a bidirectional LSTM, trained against historical approval or denial. Patent application
An LSTM is a recurrent neural network designed to carry information through a sequence. Bidirectional processing examines that available sequence in both directions. For forecasting, the entire sequence still needs to end at the prediction time; bidirectionality does not authorize access to future visits.
The design is substantive evidence that the work extends beyond isolated radiographs. A patent describes an invention and possible implementations. It does not demonstrate which components reached production, how accurately they work, or whether they improve health.
Retrace also disclosed task-specific image-quality prediction: assessing how corruption affects a particular dental task. That is a sensible engineering question. An image can be adequate for one purpose and inadequate for another. The disclosure supplies an approach to investigate, rather than a clinically validated threshold for withholding an answer. Image-quality patent
The current payer-facing site describes Retrace Connect for transactions and related workflows, and Retrace Optics for analytics. Those descriptions establish the company's stated product focus. They do not independently verify a unified longitudinal foundation model or superiority over competing products. Retrace payer products
The strongest case for this direction is practical. Correctly joining the patient's images, charting, visits and supporting documents could prevent errors that an excellent image classifier cannot address. Finding a missing attachment before submission could save rework. Separating clinical findings from policy checks could make the process easier to audit.
The same integration creates risks. Records can be linked to the wrong encounter. Care outside the network can be missing. Historical payer decisions can encode exclusions and administrative habits. A system can become very good at forecasting payment while telling us little about which treatment benefits the patient.
For that reason, I would evaluate clinical findings, prognosis and reimbursement as separate outputs, with separate reference standards. A longitudinal patient record provides useful context. Demonstrating that a treatment changes an outcome requires an additional study design that addresses who received treatment, why, and what happened afterward.

Figure 2. Parallel development paths share the same evidence requirements. Context-rich systems add record linkage, time boundaries, component testing and source traceability; each material update returns to validation.
The final model includes the people using it
A strong build should progress from locked retrospective tests to a prospective silent run, then to controlled evaluation of actual clinical use. Measure correction time, alert burden, missed findings, unnecessary interventions and relevant downstream outcomes.
A randomized crossover dental reader study illustrates why. With AI support, 22 dentists improved mean proximal-caries sensitivity from 0.72 to 0.81. Invasive and non-invasive treatment decisions also increased. The study evaluated radiograph interpretation and decisions; it did not establish better patient outcomes. 2021 reader trial
The deployment also needs correct patient matching, preserved image coordinates, unsupported-input warnings, version tracking and a safe fallback. Scanner upgrades and population changes warrant monitoring. Reviewing only clinician overrides misses errors that everybody accepted. DECIDE-AI emphasizes human factors and early live evaluation; FDA's good machine-learning-practice principles similarly treat the human–AI team and ongoing monitoring as part of development. DECIDE-AI · Good machine learning practice
That is where I would put the engineering effort: a precise target, defensible labels, strong baseline comparisons, independent testing and a workflow that can reveal failure.
Retrace's published research and disclosed designs make it a useful case to study. Public evidence reviewed here leaves substantial questions about current deployed versions, external prospective validation and patient outcomes. Those questions remain open even when the architectural idea is good.
For the next technical conversation, ask someone to walk one prediction all the way through. Which record entered? What was observed, inferred or generated? What exactly was predicted? Which unfamiliar patients tested it? What did the team do when it was wrong?
A model worth trusting should survive that walk.
Which dental AI product should we examine next?
— Thad
Source note: Public research, patent disclosures and product pages were reviewed October 4, 2026. Research results, company descriptions and editorial recommendations are distinguished above. This issue makes no claim of comparative product superiority. The prevalence example is hypothetical.
Before you buy dental AI
Download the one-page Dental AI Buyer Worksheet for your next vendor conversation. Use it to capture the evidence behind the claims and what you still need to see.
Already subscribe? Get the worksheet directly below. No need to sign up again.
Email attachment links expire after seven days. Use “Read online” at the top of this email for current download links.
New to Practice Ledger? Subscribe for free and use the one-page Dental AI Buyer Worksheet above in your next vendor conversation. It gives you a place to record the product’s job, the evidence behind it, and the questions still unanswered.
