Serum 25-hydroxyvitamin D variability in clinical trials
Every clinical trial on vitamin D begins with a measurement. Most of them end with a measurement too — and a confident claim that supplementation "raised serum levels" or that deficient participants…

Serum 25-Hydroxyvitamin D Variability in Clinical Trials: Why Your "Vitamin D Level" Is Probably Noise
Every clinical trial on vitamin D begins with a measurement. Most of them end with a measurement too — and a confident claim that supplementation "raised serum levels" or that deficient participants showed "improved outcomes." The biochemical reality is less flattering: a disturbingly large share of those numbers were never measured with the rigor the question deserves. If you are designing, auditing, or interpreting a trial that hinges on serum 25-hydroxyvitamin D [25(OH)D], the assay you choose is not a footnote. It is the experiment.
A single 25(OH)D value pulled from an unstandardized immunoassay is not a data point — it is a confession of methodological compromise.
The Gold Standard: LC-MS/MS vs. Automated Immunoassays
Let's dispose of a comfortable myth. Automated immunoassays — the workhorses of high-throughput hospital labs — are not interchangeable with liquid chromatography-tandem mass spectrometry (LC-MS/MS). They are not "close enough." They are not "good for screening." They are a different instrument answering a different question, with a different error structure.
The mechanism matters. Immunoassays rely on antibodies that bind 25(OH)D — both the D3 and D2 forms, plus a rotating cast of cross-reactants that includes vitamin D-binding protein (VDBP)-bound metabolites and, depending on the antibody clone, the C-3 epimer. When VDBP concentrations shift — as they do across ethnic groups, pregnancy, liver disease, and oral contraceptive use — recovery varies. The same serum sample can produce a different result on the same analyzer depending on what else is in the tube.
LC-MS/MS separates molecules by mass and fragmentation pattern after chromatographic resolution. It does not care about VDBP. It does not care about epimers unless the chromatographic method is poorly designed. It gives you a signal that is, with proper internal standards and calibration, traceable to a primary reference measurement procedure.
| Parameter | Automated Immunoassay | LC-MS/MS (with VDSP-compliant calibration) |
|---|---|---|
| Analytical specificity | Variable; antibody cross-reactivity with VDBP-bound metabolites and epimers | High; chromatographic separation isolates 25(OH)D3 and 25(OH)D2 |
| Typical measuring range | ~7.5–175 nmol/L | LOQ as low as 4.61 nmol/L (D3) and 1.46 nmol/L (D2) |
| Sensitivity to VDBP changes | Yes — variable recovery across populations | Negligible once internal standard corrects for ionization |
| Suitability for trial baseline/endline | Defensible only with extensive bridging and quality controls | Reference method when properly implemented |
| Bias vs. RMP | Often >5% without calibration | ≤1.7% when calibrated against certified reference material |
The uncomfortable implication: trials that ran their baseline and longitudinal samples on a single immunoassay platform — without bridging to an LC-MS/MS reference — are measuring a composite of true status and platform drift. When those numbers land in a meta-analysis, the noise compounds.
Harmonizing Trial Data with the Vitamin D Standardization Program
The Vitamin D Standardization Program (VDSP), an initiative coordinated through the NIH Office of Dietary Supplements, exists because the field recognized this exact problem. The program's reference measurement procedures, traceable to certified reference materials such as NIST SRM 2972, allow both prospective and retrospective calibration of 25(OH)D data.
This is not academic housekeeping. A 2010 analysis published in Cancer Epidemiology, Biomarkers & Prevention (CEBP) demonstrated exactly how much unstandardized immunoassay data can mislead a trial: calibration of stored serum samples shifted measured population means substantially and reclassified a non-trivial fraction of participants across deficiency thresholds. The implication was not subtle. Several high-profile observational findings about cancer and 25(OH)D looked different once the numbers were harmonized.
For a trialist, the practical checklist is short:
1. Decide your analytical platform before you enroll the first participant. Retrofitting standardization onto a completed trial is possible but expensive, and depends on the availability of stored serum with sufficient volume.
2. If you use an immunoassay, run a bridging study. Analyze at least 100–150 representative samples in parallel on your immunoassay and on a VDSP-compliant LC-MS/MS method. Apply the resulting regression to your trial data.
3. Use NIST SRM 2972 or an equivalent certified reference material to anchor your calibrators. Traceability is not optional if you want your endpoint to survive peer review in 2025 and beyond.
4. Document your chain of traceability in the methods section. Reviewers — and downstream meta-analysts — will ask. If you cannot describe the path from your reported value to a reference measurement procedure, the value is not yet data.
Addressing Isobaric Interference: The 3-epi-25(OH)D3 Challenge
Here is where the sharp-eyed laboratory scientist earns their keep. 3-epi-25(OH)D3 is a structural C-3 epimer of 25(OH)D3. Same molecular formula. Same nominal mass. Same fragmentation pattern in many MS transitions. Without chromatographic separation, it co-elutes with the analyte of interest and inflates the measured value.
This is not a theoretical concern. The epimer is present in measurable concentrations in adult serum, and at substantially higher relative concentrations in infant serum — which matters if your trial includes pediatric arms or maternal-neonatal endpoints. A 2016 CDC Stacks / Analytical Chemistry methodology paper documented the chromatographic conditions required: pentafluorophenylpropyl stationary phases, specific mobile-phase compositions, and carefully tuned gradients that resolve the epimer from the parent analyte within a reasonable run time.
| Interference | Source | Detection in LC-MS/MS | Mitigation |
|---|---|---|---|
| 3-epi-25(OH)D3 | C-3 epimer present in serum, elevated in infants | Co-elution with 25(OH)D3 on standard C18 columns | Pentafluorophenylpropyl columns with optimized gradient |
| Isobaric matrix components | Residual lipids, phospholipids | Elevated baseline or unexpected MRM transitions | Sample preparation (protein precipitation, SPE) and chromatographic tuning |
| VDBP-bound metabolites | High-VDBP samples (e.g., pregnancy) | Variable ionization suppression | Stable-isotope-labeled internal standard for 25(OH)D3 |
The mechanism is straightforward; the implementation is finicky. If your LC-MS/MS method cannot resolve the epimer, your "25(OH)D3" number is actually "25(OH)D3 plus whatever epimer sat in the same peak." For adult populations the bias is modest — single-digit nanomolar in most published comparisons — but for infant cohorts it can dominate the measurement.
If your chromatogram does not show baseline separation between the epimer and the parent compound, you are not measuring 25(OH)D3. You are measuring their sum.
Quantifying Longitudinal Within-Person Variability
Here is the number that should keep trialists honest: the within-person coefficient of variation for serum 25(OH)D over five years is approximately 14.9% (95% CI: 12.4–18.1%), with an intraclass correlation coefficient (ICC) of 0.71 (95% CI: 0.63–0.88).
What does that mean operationally? A participant with a "true" long-term mean of 50 nmol/L will, on a single measurement at a random time point, fall somewhere in a band roughly ±15% wide — roughly 42 to 58 nmol/L — simply from biological and seasonal fluctuation. The ICC of 0.71 tells you that 71% of the variance in a population's 25(OH)D values is between-person variance and 29% is within-person noise. That is a respectable ICC, but it is far from perfect, and it has direct consequences for sample size and endpoint definitions.
Three operational consequences:
- Baseline measurements are not a fixed trait. They are a snapshot. A participant classified as "deficient" at baseline by a single draw may sit comfortably above the threshold at their true long-term mean — and vice versa. Trials that rely on a single baseline value to stratify participants into exposure categories are stratifying on a noisy estimate of a slowly varying trait.
- Seasonality is real. A trial that enrolls across a calendar year will enroll different seasonal phases of 25(OH)D into different arms if not balanced. Month-of-draw is a covariate worth including.
- Change scores amplify noise. The arithmetic difference between two noisy measurements is noisier than either measurement alone. Trials powered to detect a 10 nmol/L change in an intervention arm need to account for the within-person CV in their variance estimates — or they will be underpowered.
A practical implication: when you read a trial that reports a "significant increase in serum 25(OH)D from 35 to 58 nmol/L after 12 months of supplementation," the change is larger than the within-person CV would predict by chance, so it is probably real. When a trial reports a 5 nmol/L shift that crosses a clinical threshold, the mechanism is more likely measurement noise than biology. The 14.9% within-person CV is the filter.
Technical Performance Specifications for Reference Measurement Procedures
If you are commissioning an LC-MS/MS method — or auditing a contract research organization's offering — the target performance specifications from reference measurement procedures (RMPs) are not aspirational. They are achievable and they have been demonstrated. A 2017 candidate RMP published in the Journal of AOAC International achieved:
- Total imprecision (CV) of 2.0% for 25(OH)D3 and 3.5% for 25(OH)D2
- Mean trueness of 100.4% for 25(OH)D3 and 100.3% for 25(OH)D2
- Limits of quantitation of 4.61 nmol/L (D3) and 1.46 nmol/L (D2)
These specifications exceed the VDSP performance targets of ≤5% total CV and ≤1.7% bias for routine clinical measurements, and they comfortably accommodate the longitudinal variability of real human serum. A method meeting these specifications can resolve changes well below the within-person CV band and can detect seasonal shifts of the magnitude a typical trial would expect.
What this means at the bench: if your lab's method cannot hit these specifications with appropriate reference material traceability, the limitation is not the analyte. It is the lab.
| Specification | VDSP Target | Demonstrated RMP Performance |
|---|---|---|
| Total CV (imprecision) | ≤ 5% | 2.0% (D3), 3.5% (D2) |
| Bias (trueness) | ≤ 1.7% | ~0.4% (D3), ~0.3% (D2) |
| LOQ | Method-dependent | 4.61 nmol/L (D3), 1.46 nmol/L (D2) |
| Calibration traceability | Required | NIST SRM 2972 or equivalent |
A Practical Decision Framework for Trialists
The biochemical evidence points to a single verdict: if your trial's primary or secondary endpoint is a change in serum 25(OH)D, the analytical method is part of the intervention. Cutting corners there is cutting corners in the experiment.
A defensible trial design integrates five moves:
- Run a VDSP-compliant LC-MS/MS method as the reference backbone, with NIST-traceable calibration and stable-isotope-labeled internal standards.
- Bridge any historical immunoassay data to the LC-MS/MS reference via a parallel analysis of at least 100 stored samples, and apply the resulting regression before pooling.
- Resolve the C-3 epimer chromatographically — pentafluorophenylpropyl columns or equivalent — particularly if the cohort includes infants, pregnant participants, or maternal-neonatal dyads.
- Account for the 14.9% within-person CV in power calculations, and consider repeated baseline measures (e.g., two or three draws across seasons) for trials where exposure classification matters.
- Pre-register your analytical protocol alongside the clinical protocol. Post-hoc method changes are not a path to credibility.
The blunt verdict from the laboratory bench: serum 25(OH)D is not a difficult analyte to measure accurately. It is, however, very easy to measure badly — and the difference between those two outcomes is the difference between a trial that informs policy and a trial that adds to the noise floor of an already noisy literature. The tools exist. The standardization infrastructure exists. The performance benchmarks exist. The remaining question is whether the trial you are reading — or designing — actually used them.