Methodology

Last Updated: 05/10/2026

This page explains where xDV DataVault's analytics come from, how we protect the people behind them, and how to read the results. It applies to every analytics view, export and AI answer on the platform.

Data sources

Our analytics draw on the Human Graph. This is the combined, consented data of people who contribute to xDV. Depending on what a subscriber has access to, it can include:

  • demographics that contributors tell us, such as age band, country and income band;
  • answers to questionnaires;
  • other sources a contributor has chosen to share with us, such as app and device usage or desktop browsing.

Subscribers see aggregated results only. They never see a named person, a single person's answers, or row-level data.

Consent

Contributors opt in. Each type of data has its own consent, and a contributor can turn any of them off at any time from their Data Rights page. When consent is withdrawn, that person's data is no longer used in new results. Saved and cached results are refreshed so they no longer include it.

Pseudonymisation and aggregation

Contributor data is pseudonymised. Names and contact details are kept apart from the data used for analytics. Pseudonymised data is still personal data under UK and EU data protection law, so we treat it that way.

What a subscriber receives is aggregated: counts, percentages and summary statistics for groups of people. We add further protections before any result is shown:

  • Minimum group size. Results for small groups are hidden (see below).
  • Differencing checks. We refuse a query if it is too close to recent queries. This stops someone comparing two results to work out facts about a few people.
  • Audit trail. We log every analytics query: who ran it, the settings used, and a fingerprint of the result. This keeps the system accountable.

Minimum group size

We never show a result for a group that is too small. A small group could make it possible to pick out an individual. The minimum group size depends on how identifying the combination of characteristics is. It is currently from 25 up to 100 people. Combinations that are too identifying are refused.

A hidden result shows as a blank, never as zero. A blank means "not shown", not "nobody".

Unweighted results and bases

Our results are unweighted. They describe the people in the sample, as they are. We do not adjust them to match the make-up of any population.

Every result shows its base: the number of people it is based on. Read each figure against its base. When a base is small, we show a low-base warning. In comparisons, this applies when either group has fewer than 30 people.

Statistics we use

What you seeMethodWhat it tells you
Row and column percentagesShare of the row or column baseHow answers are spread within a group
Test of association in a cross-tabChi-square test of independence, with a p-valueWhether two questions look related in this sample. We warn when expected counts are too low for the test to be reliable.
Notable cellsAdjusted residualsWhich cells differ most from what you would expect if there were no relationship
A vs B: percentagesTwo-proportion z-test, with a Newcombe 95% confidence intervalWhether two groups' percentages differ, and a likely range for the difference
A vs B: averagesWelch's t-test, with a 95% confidence intervalWhether two groups' averages differ, without assuming equal spread
Many comparisons at onceBenjamini-Hochberg false discovery rate adjustmentReduces the chance of flagging a difference that is down to chance alone

Figures with a range are labelled "Estimate with confidence interval". A confidence interval describes uncertainty from sampling within our contributors. It does not cover how those contributors differ from any wider population.

AI use and labelling

Cohort Voice answers questions written in everyday language. It uses a large language model. The model only receives aggregate statistics that have already passed the protections above. It never receives individual-level data.

Every AI answer is labelled "AI-generated" and comes with a short disclosure. Our systems also mark these answers as AI-generated in a form other software can read. We check the figures in each answer against the aggregate data. If an answer fails that check, we do not show it.

AI answers can still be wrong or incomplete. Treat them as a starting point and check the underlying tables before you rely on them.

Limitations

  • Contributors choose to join xDV. People who opt in may differ from people who do not.
  • Most data is what people tell us. Stated answers can differ from what people actually do.
  • Hidden small groups and refused queries mean some breakdowns are not available.
  • Statistical tests show whether a pattern is likely to be more than chance in this sample. They do not show cause and effect.
  • With many comparisons, some differences will appear by chance, even after adjustment.
  • Which data sources you can see depends on your subscription.

What we do not claim

  • We do not claim that our results represent the population of any country, region or market.
  • We do not weight results to population figures.
  • We do not claim that our data is anonymous. It is pseudonymised, and the results you receive are aggregated.
  • We do not claim that AI answers are free from error.

If you have a question about this page, contact us through the details on our Contact page.