optiford.

Advancing the science of micronutrient policy.

Vitamin D prevalence study: data to gather before launch

Most researchers designing a vitamin D prevalence study still anchor the protocol to a single number — usually <20 ng/mL or <50 nmol/L — and assume it travels cleanly across continents. It does not.

UpdatedAugust 27, 2026
Read time21 min read
Vitamin D prevalence study: data to gather before launch

The same serum 25(OH)D concentration has been defined as deficiency, suboptimal status, and “optimal” by different bodies. Before a single blood draw is scheduled, that threshold disagreement has to be resolved on paper, not discovered after the data lock.

A credible vitamin d deficiency prevalence study protocol therefore begins before recruitment and before the assay is selected. The study needs a declared clinical definition, a sampling calendar that accounts for sunlight, a risk-stratification plan, a laboratory method that can be compared with other studies, and enough dietary and lifestyle information to explain why two people with the same serum result may have arrived there by entirely different routes.

Defining Diagnostic Thresholds and Clinical Cutoffs

The first data point to lock in is not a lab value but a definition. And here is where the consensus stops being consensus.

A global pooled analysis covering 7,947,359 participants across 81 countries, with data collected between 2000 and 2022, reported that 15.7% of individuals fell below 30 nmol/L, 47.9% fell below 50 nmol/L, and 76.6% fell below 75 nmol/L. Three cutoffs, three prevalences, same denominator. Pick the wrong one and your headline shifts by an order of magnitude.

The 2023 update of the Polish guidelines for preventing and treating vitamin D deficiency — one of the more explicitly tiered frameworks currently in use — classifies serum concentrations as follows: deficiency at <20 ng/mL (<50 nmol/L), suboptimal status at 20–30 ng/mL (50–75 nmol/L), and optimal status at 30–50 ng/mL (75–125 nmol/L). That is a three-band scheme, not a binary one. A prevalence study that collapses it into “deficient” versus “not deficient” loses the clinically interesting middle stratum and makes comparison with other datasets harder.

By contrast, the Endocrine Society’s June 2024 guideline, Vitamin D for the Prevention of Disease, moved away from threshold-based public-health statements. It advises against routine 25(OH)D screening and unindicated supplementation in the general healthy adult population. The USPSTF reached a similar position in April 2021, concluding that the evidence was insufficient to recommend screening asymptomatic adults. Both positions matter even when the study is not evaluating clinical screening: they signal that a concentration threshold should not be treated as a universally validated boundary between health and disease.

Guideline or evidence frameworkYearStatus category or positionSerum 25(OH)D threshold
Polish national guidelines2023Deficiency<20 ng/mL (<50 nmol/L)
Polish national guidelines2023Suboptimal20–30 ng/mL (50–75 nmol/L)
Polish national guidelines2023Optimal30–50 ng/mL (75–125 nmol/L)
Global pooled meta-analysis2023Severe insufficiency<30 nmol/L
Global pooled meta-analysis2023Insufficiency<50 nmol/L
Global pooled meta-analysis2023Below “sufficient”<75 nmol/L
Endocrine Society2024No routine threshold recommended for general healthy adults
USPSTF2021Evidence insufficient to recommend screening asymptomatic adults

The protocol should distinguish three different functions that are often mixed together:

1. The primary prevalence definition. This is the cutoff used for the headline estimate and sample-size calculation.

2. Secondary categories. These show how the estimate changes under alternative thresholds and preserve comparability with established national or international frameworks.

3. Clinical or analytical sensitivity analyses. These test whether the conclusions change when borderline values, assay uncertainty, or alternative categorizations are taken into account.

The primary cutoff should be declared before recruitment, not chosen after the distribution of results becomes visible. Secondary cutoffs can be prespecified as sensitivity analyses, but they should not be presented as if they were interchangeable definitions. Each threshold needs its own unit, reference framework, and interpretation.

That distinction is particularly important when the study concerns food fortification. A fortification program may shift the population distribution upward without moving every participant above one chosen clinical boundary. If the analysis reports only the proportion below a single cutoff, it may miss a meaningful reduction in the lower tail or a change in the size of the intermediate group.

A single prevalence figure is a confession about which cutoff you chose to publish. Pick the cutoff before you choose the assay, not after.

The protocol should also state how results will be handled when they sit close to a boundary. A concentration just below 50 nmol/L is not biologically different from one just above it; the classification is an analytical and clinical convention. That does not make the cutoff useless, but it does mean that the study should report the continuous distribution as well as the categorized prevalence. Median or geometric summaries, percentile ranges, and the number of observations around each decision point can reveal whether a finding is robust or is being driven by a narrow band of borderline results.

Accounting for Seasonal Variation in Serum 25(OH)D

The second variable that breaks naive prevalence estimates is season. The same pooled analysis found that vitamin D deficiency prevalence during winter–spring was approximately 1.7 times higher than during summer–autumn. That is not a marginal effect; it is a structural one. A study that samples only in October, or only in March, produces a prevalence number that is half physiology and half calendar.

The underlying mechanism is cutaneous cholecalciferol synthesis. At latitudes above roughly 35°N, UVB radiation may not drive meaningful previtamin D3 production in the skin for several months of the year. The exact pattern depends on latitude, altitude, ozone, skin pigmentation, clothing, time spent outdoors, and behavior. Those factors do not all need to be measured with equal granularity, but the protocol should identify which ones are relevant to the target population and which are likely to distort comparisons between sites.

For protocol design, this means several practical commitments:

  • Record the calendar month and date of every blood draw, not merely the study year.
  • Define the seasonal windows before analysis. “Winter” should not mean one set of months at a northern site and another at a southern site unless that difference is part of the design.
  • Distribute recruitment across the calendar year or state clearly that the study is intentionally seasonal.
  • Report prevalence by season as well as in the annual pooled estimate.
  • Record latitude and, where relevant, altitude for each recruitment site.
  • Avoid comparing sites collected in different seasons without adjustment, stratification, or a sampling design that makes the comparison defensible.
  • Keep the lag between questionnaire completion and blood collection short enough that recent sun exposure and supplement use remain interpretable.

A common protocol error is to assume seasonality cancels out over a long enrollment period. It does not, because recruitment is rarely uniform across the calendar year. People enroll in waves that correlate with holidays, school schedules, clinical appointments, public-health campaigns, and weather. Clinic access patterns can make winter sampling easier in one region and summer sampling easier in another. The monthly distribution of blood draws should be reviewed during recruitment, not after the database is closed.

The seasonal plan also affects a fortification study’s baseline. If the intervention or policy is evaluated using samples collected before and after a fortification change, the pre-policy and post-policy samples need comparable collection windows. Otherwise, an apparent improvement may reflect the calendar rather than the food supply. The reverse is also possible: a real improvement can be obscured if the follow-up sample is collected during a period of lower endogenous synthesis.

What to record around the blood draw

The timing of the sample is only one part of seasonal control. The baseline form should capture recent conditions that can change serum 25(OH)D without representing the usual exposure pattern:

  • Time spent outdoors during the preceding weeks.
  • Travel to a lower or higher latitude.
  • Sun-protective clothing and sunscreen use.
  • A recent change in supplement use.
  • Acute illness or a period of reduced mobility.
  • Whether the participant’s work or school schedule changes by season.

These variables should not be used to manufacture a perfect exposure model. Their purpose is more modest and more useful: to identify whether a seasonal difference is plausible, whether a site has recruited an unusual group, and whether a sensitivity analysis is needed.

If your prevalence does not move between January and July, you are probably not measuring 25(OH)D — you are measuring assay noise.

Stratifying Data by Demographic and Geographical Risk

Vitamin D status is not a single population parameter. It is a vector of risk factors, and the protocol has to enumerate them before enrollment starts.

The variables that commonly stratify prevalence include:

  • Age. Older adults may have reduced cutaneous synthesis, less outdoor exposure, and a higher burden of conditions or medications that affect vitamin D metabolism. Institutionalized populations should not be treated as interchangeable with community-dwelling older adults.
  • Skin pigmentation. Higher melanin content reduces UVB-driven cholecalciferol production per unit of exposure. Skin type should be recorded with a standardized instrument or clearly defined categories rather than an improvised visual label.
  • Latitude and altitude. These shape the available UVB environment, particularly in winter, but they do not operate independently of behavior and clothing.
  • Clothing and cultural practices. The amount of skin covered can materially alter cutaneous exposure, independently of latitude.
  • Body composition. Higher adiposity is associated with lower circulating 25(OH)D through mechanisms that include sequestration and volumetric dilution. BMI may be available in most datasets, but it is not a complete substitute for more informative body-composition measures where those are feasible.
  • Pregnancy and lactation. These states alter nutritional requirements and should be analyzed explicitly rather than buried in a general adult category.
  • Malabsorption. Inflammatory bowel disease, celiac disease, and post-bariatric surgery anatomy can affect absorption and should be captured as clinical risk factors.
  • Renal and hepatic function. These are relevant to vitamin D metabolism and may change the interpretation of status in clinical subgroups.
  • Medication use. Anticonvulsants, glucocorticoids, and other therapies can influence vitamin D metabolism or the likelihood of supplementation.
  • Housing and occupation. Indoor work, institutional residence, and limited access to outdoor space can create differences within the same age and geographic group.

For each variable, the protocol must specify whether it will be used as a sampling stratum, an analytic covariate, or both. Conflating the two is a common design mistake: stratifying on age but only adjusting for BMI, or stratifying on latitude while failing to record skin type and clothing practices. Those omissions leave the analysis underpowered for the interactions that may actually drive between-group differences.

Sampling strata should be limited to variables that the study can support. If every subgroup is treated as a formal stratum, recruitment becomes fragmented and many cells become too small for a stable estimate. A more practical design may use broad age or geographic bands for recruitment, then model additional factors in the analysis. The choice should be made before enrollment and reflected in the sample-size calculation.

Geography is more than a country label

Geographical resolution matters. Reporting prevalence at the country level can hide within-country variation that exceeds between-country variation, particularly in large states with both Mediterranean and subarctic populations or in countries with substantial immigrant populations. A national estimate may be useful for policy, but it should not be mistaken for a uniform biological profile.

The site-level dataset should therefore include:

  • Geographic coordinates or a predefined latitude band.
  • Urban, peri-urban, and rural classification where relevant.
  • Altitude when it plausibly affects UV exposure.
  • The local fortification environment, including whether fortification is mandatory, voluntary, or absent for the food categories under study.
  • Recruitment setting, such as primary care, hospital, community center, workplace, school, or residential institution.
  • The calendar distribution of sampling at each site.

For an at-risk populations vitamin D assessment, recruitment setting is especially important. A hospital sample can over-represent people with chronic disease; a workplace sample can exclude people with limited employment; a household survey may under-represent institutionalized adults. These are not merely statistical nuisances. They define which population the final prevalence estimate can legitimately describe.

Standardizing Laboratory Assay Protocols for Global Comparability

Here is where the laboratory scientist earns the title. Prevalence studies live or die on assay comparability, and 25(OH)D measurement is famously difficult to standardize.

The two dominant methodological families — automated immunoassays, including chemiluminescence methods, ELISA, and RIA, and chromatographic methods such as LC-MS/MS — can disagree systematically. Immunoassays are designed to measure 25(OH)D but may show method-dependent differences in recovery of 25(OH)D2 and 25(OH)D3, as well as variable cross-reactivity with the C3-epimer and other hydroxylated metabolites. LC-MS/MS can separate and quantify the principal forms more specifically, but it requires rigorous calibration, suitable internal standards, and technical capacity that is not available in every laboratory.

The protocol must be chemically precise about what is being reported. Total 25(OH)D should be reserved for the 25-hydroxyvitamin D forms, principally 25(OH)D2 and 25(OH)D3, when those forms are being combined into one result. The C3-epimer may be reported separately where the method can resolve it and where it is relevant to the population, particularly in settings such as infancy. 24,25-dihydroxyvitamin D is a distinct metabolite and is not part of total 25(OH)D; it should be reported separately. It must not be added to a total 25(OH)D value or described as one of its component forms.

That distinction is not a matter of wording. A method may measure 24,25-dihydroxyvitamin D as part of a broader vitamin D metabolite panel, but its result belongs in a separate variable. Combining it with 25(OH)D would change the analyte definition and make comparison with standard prevalence thresholds invalid.

For a study intended to produce data comparable with the international literature, the laboratory section should settle at least the following:

  • Analytical platform and instrument model.
  • Manufacturer, reagent formulation, and lot-tracking procedure.
  • Whether 25(OH)D2 and 25(OH)D3 are measured separately or reported as a combined total.
  • Whether the C3-epimer is resolved, reported separately, or excluded from the reported 25(OH)D result according to the method’s validated definition.
  • Whether 24,25-dihydroxyvitamin D is measured; if it is, specify that it will be stored and reported as a separate metabolite.
  • Analytical measuring range, limit of detection, and limit of quantification.
  • Precision estimates, including the coefficient of variation at concentrations close to the study’s clinical decision points.
  • Calibration and traceability to an appropriate reference measurement procedure.
  • Sample type, processing time, storage temperature, freeze–thaw limits, and shipment conditions.
  • Rules for repeat testing, dilution, hemolysis, insufficient volume, and results below quantification.

Participation in an external quality-assurance scheme such as DEQAS, the Vitamin D External Quality Assessment Scheme, or the CDC’s Vitamin D Standardization-Certification Program can support alignment with reference procedures. The protocol should not simply name a scheme. It should state how proficiency results will be reviewed, how drift will be handled, and whether stored samples will be reanalyzed if a calibration problem is identified.

One platform is not automatically one result

If the study is multi-site, the protocol must specify whether samples will be analyzed centrally or locally. Local analysis introduces between-site variance that can exceed the within-site variance the study is trying to detect. Central analysis on a single platform is usually the cleaner design, provided sample stability, transport, and laboratory capacity are adequate.

When local testing is unavoidable, a bridging study should be planned before the main analysis. A shared set of samples can be measured across participating laboratories, allowing the investigators to estimate systematic differences rather than applying a generic conversion factor borrowed from another population or assay generation. Any correction model must be prespecified, retained with the analytical dataset, and accompanied by uncertainty estimates.

The unit policy also needs to be explicit. Concentrations may be reported in ng/mL or nmol/L, and the conversion is commonly expressed as 1 ng/mL = 2.5 nmol/L. That conversion changes units; it does not correct for assay bias. A mathematically correct unit conversion cannot make two non-equivalent assays comparable.

The single most expensive mistake in vitamin D epidemiology is to discover, after publication, that two labs used two assays and a conversion factor. Build the assay plan before you build the recruitment plan.

The assay plan should also distinguish the primary endpoint from exploratory metabolite data. If the study’s prevalence endpoint is based on serum total 25(OH)D, then additional measurements such as 24,25-dihydroxyvitamin D should not silently redefine that endpoint. They may provide valuable information about metabolism, renal handling, or response to fortification, but they answer a different question.

Integrating Baseline Dietary and Lifestyle Exposure Variables

Serum 25(OH)D is an integrated biomarker. It reflects cutaneous synthesis over the preceding weeks, recent oral intake from food and supplements, adiposity, and hydroxylation capacity. Without the behavioral and dietary context, the serum value is difficult to interpret. A high summer reading and a high supplement-related reading may look identical in the laboratory but mean very different things for public-health planning.

The minimum baseline dataset should include:

  • Self-reported frequency of oral vitamin D supplementation, with dose, formulation, and duration if possible.
  • Dietary intake assessment focused on fortified food categories, including dairy products, margarine, breakfast cereals, plant-based alternatives, and infant formula where relevant.
  • Sun-exposure patterns, including time outdoors on working and non-working days.
  • Clothing coverage during periods of peak UV exposure and sunscreen use.
  • Recent travel to lower latitudes, particularly during the months immediately preceding blood collection.
  • Skin-type classification using a standardized scale such as Fitzpatrick I–VI or an equivalent instrument, with self-report clearly identified.
  • Relevant medical conditions, medications, pregnancy or lactation status, and recent changes in mobility or outdoor activity.

This is precisely the data layer that the European Commission’s OPTIFORD project, launched in 2001 under the Fifth Framework Programme to investigate food fortification as a strategy for insufficient vitamin D status, was designed to help develop across European populations. A fortification study cannot evaluate the food environment while treating individual intake as a black box.

A common design shortcut is to skip detailed dietary assessment because national fortification policies are assumed to homogenize intake. They do not. Voluntary fortification means that intake within a single country can vary substantially with brand choice, household income, dietary pattern, and whether a person consumes the relevant food category at all. A national policy describes the market; it does not tell you what any particular participant ate.

The 800 IU, or 20 µg, daily D-A-CH reference intake for adults without endogenous synthesis is a population target, not an observed average. It should not be used as a substitute for measured intake or as evidence that participants received that amount. In an analysis of food fortification, the distinction between recommended intake, estimated intake, and actual product consumption needs to remain visible.

Match the instrument to the question

The protocol should name the instrument used for each exposure and define its recall window:

  • A food-frequency questionnaire can describe usual patterns but may be weak at estimating the amount of a newly fortified product consumed.
  • A 24-hour recall provides more detailed intake data but may not represent usual behavior from one day.
  • A short dietary diary can capture product-level information, although completion burden may affect participation and missingness.
  • A brief sun-exposure diary may be practical during a seasonal substudy, while a retrospective questionnaire may be more suitable for a large population survey.

Brand and product information can be important when fortification levels vary across the market. The study does not need to turn every participant into a dietary chemist, but it should collect enough detail to distinguish fortified from non-fortified products and to identify the main sources of vitamin D. If nutrient databases are used, the version and assumptions behind the database should be documented.

Missing exposure data should be handled as part of the design, not as an afterthought. The analysis plan should state which variables are essential for the primary model, how “unknown” responses will be coded, when multiple imputation is appropriate, and whether a complete-case analysis will be used as a sensitivity analysis. Outlier rules also need to be set before investigators see the association between intake and serum concentration.

Turning the protocol into an auditable study

Before the first participant is enrolled, the protocol should be capable of answering a simple question: could another research team reproduce the measurement and classification decisions without asking the investigators what they meant?

That requires a written record of the following decisions:

1. The primary and secondary 25(OH)D cutoffs, with the reference framework for each and a sensitivity-analysis plan.

2. The recruitment calendar, including monthly targets or a documented reason for intentionally seasonal sampling.

3. The demographic and geographical variables that define sampling strata, analytic covariates, or both.

4. The laboratory platform, calibration procedure, external quality-assurance participation, and reagent-lot tracking.

5. The definition of reported total 25(OH)D as a measure of the 25-hydroxyvitamin D forms, with 24,25-dihydroxyvitamin D kept as a separate metabolite whenever it is measured.

6. The central-versus-local analysis decision and the bridging procedure for a multi-site study.

7. The minimum dietary and lifestyle dataset, including supplements, fortified foods, sun exposure, travel, and skin type.

8. The rules for sample processing, storage, repeat testing, unit conversion, missing data, and outliers.

9. The denominator for every prevalence estimate, including exclusions and handling of participants with incomplete laboratory or questionnaire data.

These are not administrative details surrounding the science. They are the science. If the study cannot explain why one participant was classified as deficient, why another was placed in an intermediate category, and whether both results were produced under the same analytical definition, the final percentage will be more precise-looking than reliable.

The audit should be repeated during recruitment. Monthly monitoring can identify a site that is collecting almost exclusively in one season, enrolling too few participants from a key risk group, or sending samples to a different laboratory than the one named in the protocol. Correcting those problems while enrollment is open is far less damaging than trying to reconstruct the sampling frame after publication.

What the final prevalence estimate can actually say

A vitamin D prevalence estimate is only as generalizable as its sampling frame, and only as comparable as its analyte definition. A community sample of adults cannot automatically describe nursing-home residents, pregnant women, children, or people with malabsorption. A winter sample cannot be treated as an annual estimate without a defensible seasonal design. A result from one immunoassay cannot be compared with an LC-MS/MS result simply because both are labeled serum 25(OH)D.

This is also why the baseline dataset matters for food fortification. If serum status improves after a fortification policy is introduced, the interpretation depends on whether supplement use changed, whether the composition of the sampled population shifted, whether the follow-up season was comparable, and whether the laboratory method remained stable. Without those controls, the study may observe a change but remain unable to attribute it.

The most useful paper may not be the one with the largest headline percentage. It may be the one that reports the continuous distribution, shows the result under several prespecified thresholds, separates seasons and risk groups, identifies the assay method in enough detail for comparison, and makes clear which findings apply to the sampled population rather than to an entire country or continent.

The design has to come before the number

A vitamin D prevalence study is not a blood test repeated at scale. It is a measurement exercise whose result is determined more by protocol decisions — which cutoff, which season, which assay, which covariates, and which population — than by the population being sampled. The biochemistry is the easy part. The hard part is locking the design before anyone sees a single nmol/L.

Get the design wrong, and the literature gets another prevalence figure that cannot be compared with anything published in the last five years. Get it right, and the data can contribute to the global picture — the same picture that currently shows 47.9% of people in 81 countries below 50 nmol/L, a number whose meaning depends entirely on what the investigators decided that cutoff represented before the study began.

FAQ

Which vitamin D threshold should a prevalence study use?
The study should declare a primary cutoff before recruitment and identify the reference framework behind it. Secondary thresholds can be prespecified for sensitivity analyses, but they should not be treated as interchangeable definitions.
How does season affect vitamin D deficiency prevalence?
A global pooled analysis reported that deficiency prevalence during winter–spring was approximately 1.7 times higher than during summer–autumn. Studies should record the date and month of each blood draw, define seasonal windows in advance, and report prevalence by season.
What participant information should be collected in a vitamin D prevalence study?
The baseline dataset should include supplement use, intake of fortified foods, sun exposure, clothing coverage, sunscreen use, recent travel, skin type, relevant medical conditions, medications, pregnancy or lactation status, and changes in mobility or outdoor activity.
Can results from different vitamin D assays be compared directly?
Not necessarily. Immunoassays and LC-MS/MS can differ systematically, so multi-site studies should use central analysis where feasible or conduct a prespecified bridging study with shared samples when local testing is unavoidable.
Is 24,25-dihydroxyvitamin D part of total 25(OH)D?
No. Total 25(OH)D refers to the principal 25-hydroxyvitamin D forms, mainly 25(OH)D2 and 25(OH)D3. 24,25-dihydroxyvitamin D is a distinct metabolite and should be reported separately.