optiford.

Advancing the science of micronutrient policy.

Serum 25(OH)D assays: avoiding costly trial errors

Clinical vitamin D trials can fail at the measurement stage before any biological conclusion is reached. The problem is not limited to an incorrect instrument reading.

UpdatedAugust 28, 2026
Read time16 min read
Serum 25(OH)D assays: avoiding costly trial errors

A systematic assay bias can misclassify baseline vitamin D status, distort the apparent treatment effect, and obscure the difference between vitamin D2 and vitamin D3 interventions.

[Serum 25-hydroxyvitamin D [25(OH)D] is the standard](/articles/vitamin-d-tests-how-seasonal/) marker used to assess vitamin D status. Its analytical reliability therefore controls the interpretation of deficiency prevalence, supplementation response, fortification impact, and associations with bone or immune outcomes. If measurements are not traceable to a reference procedure, the trial may compare analytical variation rather than physiology.

The operational requirement is defined by the Vitamin D Standardization Program (VDSP). For 25(OH)D assays, the target is imprecision of no more than 10% and bias of no more than 5% relative to isotope dilution liquid chromatography-tandem mass spectrometry (ID LC-MS/MS) reference values. These are not administrative specifications. They determine whether a concentration difference represents a true intervention signal.

The hidden cost of assay variability in clinical research

A vitamin D trial usually assumes that a reported serum concentration is a stable quantitative observation. In practice, the result is produced by a chain of pre-analytical, analytical, and calibration decisions.

The relevant chain includes:

1. Sample collection and handling. Hemolysis, icterus, and lipemia can interfere with routine immunoassays. The effect is not uniform across platforms.

2. Analyte recognition. The assay must detect both 25-hydroxyvitamin D3 and, where relevant, 25-hydroxyvitamin D2.

3. Calibration. The result must be aligned with reference target values rather than only with the manufacturer’s internal calibrator system.

4. Specificity control. The method must distinguish 25(OH)D from structurally related compounds, including the 3-epi-25-hydroxyvitamin D3 epimer.

5. Precision across the study run. Replicate measurements must remain sufficiently close to each other to prevent random noise from expanding confidence intervals.

A bias at baseline changes group classification. A bias during follow-up changes the estimated response. Random imprecision weakens statistical power. These mechanisms are different and require different controls.

A trial can therefore have an apparently adequate sample size and still produce weak evidence if the measurement system is poorly standardized. The issue is particularly serious in studies using threshold categories, such as deficiency or sufficiency groups. A modest analytical shift can move participants across a category boundary without any change in underlying vitamin D physiology.

A 25(OH)D value is not automatically comparable to another 25(OH)D value. The assay platform, calibration system, analyte composition, and interference profile define the measurement.

The VDSP Intercomparison Study 2 illustrates the scale of the problem. Among 32 evaluated ligand binding assays used across 17 laboratories, only 50% met the VDSP criterion for mean bias of no more than ±5% relative to reference target values. This finding does not mean that every non-compliant assay produces unusable data in every context. It means that platform performance cannot be assumed from the assay label alone.

For investigators, the practical consequence is direct: the laboratory method belongs in the trial protocol as a controlled study variable. It should not be treated as an invisible technical detail.

The VDSP framework separates two performance dimensions: bias and imprecision.

Bias is the systematic difference between the measured value and the reference value. If an assay consistently reads low, the trial may report a lower vitamin D status than actually present. If it reads high, deficiency may be under-reported.

Imprecision is the random variation of repeated measurements. An imprecise assay may produce a different result when the same sample is measured again. This variation reduces the ability to detect a treatment-related change.

The standard VDSP performance criteria are:

ParameterVDSP targetConsequence of failure
Mean bias relative to ID LC-MS/MS reference values≤5%Systematic overestimation or underestimation of vitamin D status
Imprecision≤10%Excessive random variation between measurements
Retrospective calibration input40–50 single-donor banked serum samplesInsufficient basis for a reliable historical correction equation
Intercomparison result in Study 250% of evaluated ligand binding assays met the bias criterionPlatform suitability must be demonstrated, not presumed

The reference method is isotope dilution LC-MS/MS. This designation is important, but it does not make every LC-MS/MS result automatically correct. Mass spectrometry still requires appropriate chromatographic separation, calibration, quality control, and analyte definition. In particular, C-3 epimers can create an analytical problem if they are not separated before quantification.

For a new clinical trial, the laboratory should provide evidence that its method is standardized against reference values. The evidence should cover the actual platform, reagent lot conditions, calibration procedure, and sample matrix used in the study. A general statement that a laboratory performs vitamin D testing is not sufficient technical documentation.

The study team should also predefine how results will be handled when samples are compromised. Hemolyzed, icteric, or lipemic specimens should not be silently combined with clean samples if the platform is known to respond differently to these conditions. At minimum, the protocol should specify whether such samples are excluded, repeated, flagged, or analyzed in a sensitivity analysis.

Why threshold-based outcomes are especially vulnerable

Many vitamin D studies use categorical endpoints. Examples include the proportion of participants below a prespecified serum 25(OH)D concentration or the percentage reaching a target after intervention. These endpoints amplify small analytical shifts.

Consider two groups with values concentrated near the decision threshold. A systematic assay bias can change the number classified as deficient without changing the distribution of actual concentrations. If the same method is used at baseline and follow-up, some systematic error may cancel when the analysis focuses on within-person change. It will not necessarily cancel when the endpoint is a prevalence estimate, a between-study comparison, or a subgroup classification.

The risk increases when different laboratories or assay platforms are used for different trial phases. A change in platform can appear as a biological response. This is a form of measurement discontinuity and should be handled as a protocol deviation or calibration problem, not interpreted as treatment efficacy.

Immunoassay limitations: D2 under-recognition and sample interference

Automated immunoassays remain common because they offer high throughput and integration with routine clinical laboratory systems. Their limitations become more consequential when the intervention contains vitamin D2 or when the study includes samples with substantial matrix variation.

Vitamin D2 is not a minor technical detail

Many automated immunoassays under-recognize or under-estimate 25-hydroxyvitamin D2 relative to 25-hydroxyvitamin D3. This creates a specific problem for trials involving vitamin D2 supplementation or food fortification with ergocalciferol.

A study can administer vitamin D2 successfully and still show a weaker serum response than expected if the assay detects the metabolite inefficiently. The apparent result may be interpreted as low bioavailability, poor adherence, inadequate dose, or limited physiological conversion. The analytical explanation may be simpler: the method has unequal recovery for the two major forms.

This is why assay selection must follow the intervention chemistry. The relevant question is not only whether the platform measures total 25(OH)D. It is whether it measures the metabolites generated by the specific vitamin D compound used in the trial.

For a D2 intervention, the laboratory should document:

  • performance data for 25OH D2 and 25OH D3 separately;
  • calibration traceability for both metabolites;
  • whether the method reports total 25(OH)D or individual metabolite concentrations;
  • expected cross-reactivity or recovery differences;
  • compatibility with the study’s target population and sample matrix.

A D3-focused assay validation cannot be transferred automatically to a D2 fortification study.

Matrix effects in routine samples

Serum is a variable biological matrix. Lipid content, bilirubin, hemoglobin, protein composition, and other sample characteristics can affect assay behavior. Routine clinical immunoassays exhibit significant analytical interferences in hemolyzed, icteric, or lipemic samples. The measured 25(OH)D concentration can differ from a mass spectrometry result obtained from the same specimen.

This does not mean that every visibly altered sample is invalid. It means that the laboratory must know the platform-specific interference limits and apply them consistently. A trial that records sample quality but does not connect quality flags to analytical review loses an opportunity to control a known source of variation.

The most defensible workflow is operational:

1. Record the sample condition using the laboratory’s established indices.

2. Apply platform-specific acceptance or rejection limits.

3. Repeat or confirm results when the concentration is clinically or statistically influential.

4. Preserve the original result and the corrected or confirmatory result in the trial database.

5. Include a sensitivity analysis if the affected samples are retained.

This procedure is more useful than a generic statement that the assay was performed according to routine laboratory practice. Clinical routine and clinical research have different tolerance for undocumented analytical variation.

The epimer challenge: why LC-MS/MS requires chromatographic precision

The compound 3-epi-25-hydroxyvitamin D3, commonly abbreviated 3-epi-25OHD3, has the same nominal mass as 25(OH)D3. This creates isobaric interference in LC-MS/MS assays unless the compounds are chromatographically separated.

The consequence is overestimation of total 25(OH)D concentrations when the epimer is included in the reported signal but is not part of the intended analyte definition. The issue is particularly relevant in infant samples. The epimer can constitute up to 58% of total 25(OH)D3 in neonatal blood samples.

This is not a reason to reject mass spectrometry. It is a reason to specify the mass spectrometry method in greater detail. “LC-MS/MS” describes an analytical family, not a single performance level. The method must state whether chromatographic separation of the C-3 epimer is performed and how the result is reported.

The study documentation should distinguish among at least three possible reporting approaches:

  • 25(OH)D3 measured with epimer separation. The epimer is resolved chromatographically and excluded or reported separately.
  • Combined signal reported as 25(OH)D3. The result may be inflated if the epimer contributes materially to the signal.
  • Total 25(OH)D reported without full metabolite resolution. The interpretation depends on the assay’s defined measurand and calibration procedure.

The clinical bioactivity of 3-epi-25(OH)D3 relative to 25(OH)D3 in non-skeletal pathways remains incompletely characterized. This limits the biological interpretation of a combined measurement. A higher analytical concentration is not automatically equivalent to a higher concentration of biologically interchangeable 25(OH)D3.

LC-MS/MS versus immunoassay: the actual comparison

The common comparison, “LC-MS/MS versus immunoassay,” is too broad to support a procurement decision. The relevant comparison is between validated methods under defined conditions.

Technical featureAutomated immunoassayLC-MS/MS
ThroughputUsually high and compatible with routine analyzersVariable; depends on extraction, chromatography, and batch design
25OH D2 recognitionMay under-recognize or under-estimate relative to 25OH D3Can quantify metabolites separately when the method is validated for both
Matrix interferenceCan be significant in hemolyzed, icteric, or lipemic samplesReduced in some workflows but not eliminated
3-epi-25OHD3 controlPlatform-dependent; may not resolve the epimerRequires chromatographic separation to prevent isobaric interference
Standardization requirementMust be assessed against VDSP targetsMust also be calibrated, validated, and quality-controlled
Main operational riskUnequal antibody recognition and matrix effectsInadequate separation, calibration drift, or incorrect measurand definition

The appropriate method depends on the trial design. A high-throughput fortification survey may prioritize standardized processing across many samples. A mechanistic trial involving infants, vitamin D2, or metabolite-specific outcomes may require a method with explicit separation and recovery data. The decision is not a contest between instrument categories. It is a comparison of validated performance against the study’s biological and statistical requirements.

LC-MS/MS is a reference-oriented technology only when the measurand, calibration, chromatographic separation, and quality controls are defined.

Retrospective calibration: salvaging historical trial datasets

Historical vitamin D studies often contain valuable samples and outcome data but were conducted before current standardization procedures were applied. Repeating the clinical intervention is usually impossible. The remaining option is to determine whether the historical measurements can be linked to a reference system.

The VDSP established protocols for retrospective laboratory standardization using 40 to 50 single-donor banked serum samples. These samples are used to generate linear calibration equations linking historical trial data to certified reference procedures.

The method is not a universal correction factor. A calibration equation must reflect the original assay, laboratory, time period, sample distribution, and storage conditions. Applying a generic adjustment to results from an unrelated platform can introduce a second layer of bias.

A controlled retrospective workflow contains several stages.

1. Identify the original measurement system

The study team should establish which assay platform, reagent generation, calibration system, and laboratory performed the original measurements. If the platform changed during the trial, the data should be stratified by method and period.

An instrument name alone is insufficient. Assay versions and reagent lots can affect performance. The historical laboratory record is therefore part of the analytical dataset.

2. Locate suitable banked sera

The calibration material should represent the original sample population across the concentration range relevant to the trial. The VDSP protocol uses 40–50 single-donor banked serum samples to construct the relationship between historical results and reference values.

Pooling samples can obscure individual matrix effects and is not equivalent to a single-donor calibration set. The objective is to characterize how the original method behaves across real samples, not only across manufactured controls.

3. Measure against a reference procedure

The banked samples are analyzed using a certified reference procedure, commonly based on ID LC-MS/MS. The method must address the epimer issue where relevant. If the reference measurement itself does not separate 3-epi-25OHD3, the target value must be defined accordingly.

The historical assay results and reference values are then used to derive a calibration equation. The equation should be evaluated for residual error, concentration-dependent behavior, and plausibility at the lower and upper ends of the range.

4. Apply the correction transparently

Corrected values should not replace the original data without an audit trail. The database should preserve:

  • original reported concentration;
  • original assay and laboratory;
  • reference value used for calibration;
  • calibration equation version;
  • corrected concentration;
  • inclusion or exclusion status;
  • analytical flags and missingness codes.

This is necessary for reproducibility. It also permits secondary analyses using either the uncorrected or standardized dataset.

5. Reassess the clinical conclusions

Standardization can change deficiency prevalence, group assignment, dose-response estimates, and associations with clinical outcomes. The analysis should therefore be repeated after calibration rather than limited to a statement that the mean concentration changed.

The primary question is whether the direction, magnitude, and statistical uncertainty of the trial conclusion remain stable. If they do not, the original result was sensitive to assay scale. That is an analytical finding, not an inconvenience to be hidden.

Building a measurement plan before the first sample is collected

The lowest-cost point for controlling serum 25-hydroxyvitamin D assay errors is before recruitment. Once samples have been tested on an unsuitable or undocumented platform, correction may be possible but not guaranteed.

A practical measurement plan should define the following:

  • Target analyte. Total 25(OH)D, 25OH D2, 25OH D3, or a combination with separate reporting.
  • Intervention form. Cholecalciferol, ergocalciferol, or fortified food containing a defined vitamin D compound.
  • Reference alignment. Evidence that the laboratory method meets VDSP bias and imprecision criteria.
  • Epimer handling. Whether 3-epi-25OHD3 is separated and how it enters the final result.
  • Sample quality rules. Treatment of hemolysis, icterus, lipemia, inadequate volume, and repeated freeze-thaw exposure.
  • Inter-laboratory control. Procedures for harmonizing results if more than one laboratory is involved.
  • Longitudinal consistency. Restrictions on platform changes during the trial and bridging studies if a change is unavoidable.
  • Statistical treatment. Prespecified sensitivity analyses for assay batch, laboratory, sample quality, and corrected versus uncorrected values.

The protocol should also avoid mixing analytical decisions with post hoc clinical interpretation. For example, a concentration should not be reclassified as deficient merely because the assay result is lower than expected after supplementation. First establish whether the measurement is traceable and technically valid. Only then interpret the biological response.

A compact decision route

For most trials, the decision can be reduced to five technical questions:

1. Does the method meet the VDSP target of no more than 5% mean bias and 10% imprecision?

2. Does it measure the vitamin D metabolite generated by the intervention with acceptable recovery?

3. Are hemolyzed, icteric, and lipemic samples controlled under documented rules?

4. Is the 3-epi-25OHD3 epimer separated or otherwise accounted for?

5. Can the same measurement scale be maintained from baseline through follow-up?

A “no” answer does not always terminate the study. It identifies the required mitigation: a different platform, confirmatory testing, stratified analysis, retrospective calibration, or a narrower interpretation of the endpoint.

Cost-benefit summary of the technical choice

The economic comparison is not simply the per-sample price of immunoassay versus LC-MS/MS. The relevant cost includes repeat testing, sample exclusion, laboratory bridging, statistical reanalysis, delayed database lock, and the loss of interpretability when a treatment effect cannot be separated from assay bias.

Automated immunoassay can be efficient when the platform is standardized, the intervention is compatible with its analyte recognition profile, and sample quality is controlled. It becomes a weak choice when the study depends on accurate D2 recovery, includes a high proportion of compromised samples, or requires metabolite-specific interpretation.

LC-MS/MS provides greater analytical flexibility, particularly for separate metabolite measurement. It still requires method validation, chromatographic separation, calibration discipline, and quality control. Selecting mass spectrometry without specifying epimer separation or the target measurand does not resolve the measurement problem.

Retrospective standardization can recover value from historical datasets. It requires appropriate banked sera, reference measurements, and a calibration model specific to the original analytical system. It cannot convert undocumented measurements into fully reliable data by assumption.

The strict conclusion is therefore narrow. Serum 25(OH)D assay errors are preventable when measurement is treated as a controlled component of trial design. The minimum technical route is VDSP-aligned performance, explicit D2 and D3 recognition data, documented control of matrix interference, chromatographic consideration of 3-epi-25OHD3, and preservation of the full calibration record. The additional laboratory effort is justified when the endpoint depends on small changes in vitamin D status or on classification near a clinical threshold. Without that control, the trial may quantify assay behavior with the appearance of nutritional science.

FAQ

What are the VDSP performance targets for 25(OH)D assays?
The VDSP targets are mean bias of no more than 5% and imprecision of no more than 10% relative to ID LC-MS/MS reference values.
Why can vitamin D assays produce misleading clinical trial results?
Systematic bias can change baseline classification and estimated treatment response, while random imprecision reduces the ability to detect treatment-related changes. Small analytical shifts can also move participants across deficiency or sufficiency thresholds.
Can immunoassays accurately measure vitamin D2?
Many automated immunoassays under-recognize or under-estimate 25-hydroxyvitamin D2 relative to 25-hydroxyvitamin D3. This can make a vitamin D2 intervention appear to produce a weaker serum response than it actually does.
Why must 3-epi-25OHD3 be separated in some LC-MS/MS methods?
3-epi-25OHD3 has the same nominal mass as 25(OH)D3, creating isobaric interference unless the compounds are separated chromatographically. If the epimer contributes to the reported signal, total 25(OH)D concentrations may be overestimated.
How should hemolyzed, icteric, or lipemic samples be handled?
The laboratory should record sample condition, apply platform-specific acceptance or rejection limits, and repeat or confirm results when they are influential. If affected samples are retained, the study should include a sensitivity analysis.
Can historical vitamin D trial results be retrospectively standardized?
They may be linked to a reference system using 40–50 single-donor banked serum samples and a calibration equation specific to the original assay, laboratory, time period, sample distribution, and storage conditions. Corrected values should be preserved alongside the original data with a full audit trail.