Methodology
Last Updated: 05/10/2026
This page explains where xDV DataVault's analytics come from, how we protect the people behind them, and how to read the results. It applies to every analytics view, export and AI answer on the platform.
Data sources
Our analytics draw on the Human Graph. This is the combined, consented data of people who contribute to xDV. Depending on what a subscriber has access to, it can include:
- demographics that contributors tell us, such as age band, country and income band;
- answers to questionnaires;
- other sources a contributor has chosen to share with us, such as app and device usage or desktop browsing.
Subscribers see aggregated results only. They never see a named person, a single person's answers, or row-level data.
Consent
Contributors opt in. Each type of data has its own consent, and a contributor can turn any of them off at any time from their Data Rights page. When consent is withdrawn, that person's data is no longer used in new results. Saved and cached results are refreshed so they no longer include it.
Pseudonymisation and aggregation
Contributor data is pseudonymised. Names and contact details are kept apart from the data used for analytics. Pseudonymised data is still personal data under UK and EU data protection law, so we treat it that way.
What a subscriber receives is aggregated: counts, percentages and summary statistics for groups of people. We add further protections before any result is shown:
- Minimum group size. Results for small groups are hidden (see below).
- Differencing checks. We refuse a query if it is too close to recent queries. This stops someone comparing two results to work out facts about a few people.
- Audit trail. We log every analytics query: who ran it, the settings used, and a fingerprint of the result. This keeps the system accountable.
Minimum group size
We never show a result for a group that is too small. A small group could make it possible to pick out an individual. The minimum group size depends on how identifying the combination of characteristics is. It is currently from 25 up to 100 people. Combinations that are too identifying are refused.
A hidden result shows as a blank, never as zero. A blank means "not shown", not "nobody".
Unweighted results and bases
Our results are unweighted. They describe the people in the sample, as they are. We do not adjust them to match the make-up of any population.
Every result shows its base: the number of people it is based on. Read each figure against its base. When a base is small, we show a low-base warning. In comparisons, this applies when either group has fewer than 30 people.
Statistics we use
| What you see | Method | What it tells you |
|---|---|---|
| Row and column percentages | Share of the row or column base | How answers are spread within a group |
| Test of association in a cross-tab | Chi-square test of independence, with a p-value | Whether two questions look related in this sample. We warn when expected counts are too low for the test to be reliable. |
| Notable cells | Adjusted residuals | Which cells differ most from what you would expect if there were no relationship |
| A vs B: percentages | Two-proportion z-test, with a Newcombe 95% confidence interval | Whether two groups' percentages differ, and a likely range for the difference |
| A vs B: averages | Welch's t-test, with a 95% confidence interval | Whether two groups' averages differ, without assuming equal spread |
| Many comparisons at once | Benjamini-Hochberg false discovery rate adjustment | Reduces the chance of flagging a difference that is down to chance alone |
Figures with a range are labelled "Estimate with confidence interval". A confidence interval describes uncertainty from sampling within our contributors. It does not cover how those contributors differ from any wider population.
AI use and labelling
Cohort Voice answers questions written in everyday language. It uses a large language model. The model only receives aggregate statistics that have already passed the protections above. It never receives individual-level data.
Every AI answer is labelled "AI-generated" and comes with a short disclosure. Our systems also mark these answers as AI-generated in a form other software can read. We check the figures in each answer against the aggregate data. If an answer fails that check, we do not show it.
AI answers can still be wrong or incomplete. Treat them as a starting point and check the underlying tables before you rely on them.
Limitations
- Contributors choose to join xDV. People who opt in may differ from people who do not.
- Most data is what people tell us. Stated answers can differ from what people actually do.
- Hidden small groups and refused queries mean some breakdowns are not available.
- Statistical tests show whether a pattern is likely to be more than chance in this sample. They do not show cause and effect.
- With many comparisons, some differences will appear by chance, even after adjustment.
- Which data sources you can see depends on your subscription.
What we do not claim
- We do not claim that our results represent the population of any country, region or market.
- We do not weight results to population figures.
- We do not claim that our data is anonymous. It is pseudonymised, and the results you receive are aggregated.
- We do not claim that AI answers are free from error.
If you have a question about this page, contact us through the details on our Contact page.