Datasets
Dataset Evaluation Checklist
A practical guide to help you assess and compare sustainability datasets before integrating them into your investment workflow.
- Assets, issuers, countries, sectors, time periods covered
- Point‑in‑time history vs only latest snapshot
- Percentage of your portfolio/universe covered
- Treatment of private markets, sovereigns, derivatives
- Clear documentation of how each metric is calculated
- Alignment with major frameworks (ISSB, TCFD, SFDR, GHG Protocol, etc.)
- Handling of estimates vs reported data; use of models and assumptions
- Version control and change logs for methodology updates
- Accuracy: Cross‑checks against primary sources or filings
- Completeness: Missing values, empty fields, coverage by sector/region
- Consistency: Same definitions over time, across regions and asset classes
- Outlier handling: Rules for flags, winsorization, or overrides
Missing data rates
- % of nulls per field, per sector, per region
- Are missing values clustered (e.g., EM, small caps, private assets)?
Coverage vs portfolio
- % of AUM with non‑missing values for each key metric
- List of top 20 holdings without data
Distribution sanity checks
- Min / max / mean / median for each metric
- Compare to expected ranges (e.g., CO₂ intensity)
- Look for impossible values (negative where impossible, >100% where not allowed)
Outlier detection
- Z‑scores or percentile bands (e.g., flag values beyond 3σ or 1st/99th percentile)
- Check whether outliers are data errors or genuine edge cases
Time‑series stability
- Year‑on‑year changes; flag jumps above a threshold (e.g., >50% change)
- Count of restatements or backfills over time
Cross‑metric consistency
- Ratios that should reconcile (e.g., Scope 1+2 vs total emissions)
- Check that derived metrics match their components when you recompute them
Duplicate & uniqueness checks
- One record per issuer/ISIN/date where expected
- No duplicated rows or conflicting values for the same key
Aggregation checks
- Aggregate security‑level data and compare with any portfolio‑level numbers provided
- Recompute vendor‑supplied scores from raw fields on a sample to confirm formulas
- How often data is refreshed (daily / monthly / annually)
- Typical lag between real‑world event/report and dataset update
- Backfill policies when new information becomes available
- Ability to drill down to source documents or raw inputs
- Clear audit trail (who changed what, when, and why)
- Availability of data dictionaries and field‑level metadata
- Coverage of metrics needed for CSRD/SFDR/Taxonomy, SDR, SEC, ISSB, etc.
- Mapping tables from dataset fields to reporting templates or PAIs
- Vendor roadmap for keeping up with new regulations
- Delivery options: API, SFTP, bulk CSV/Parquet, Excel add‑ins
- Identifier coverage and mapping (ISIN, LEI, ticker, internal IDs)
- Performance: file sizes, latency, rate limits, and support for backfills
- Track record, client base similar to you (AM, pension, insurance)
- SLAs for uptime, bug fixes, and data corrections
- Quality of documentation, user guides, and client success support
- Clear support channels in case of missing files or API disruptions
- Access to knowledgeable contacts for questions regarding the data, methodology, and underlying assumptions
- Does the granularity match your needs (security‑level vs scores only)?
- Suitability for risk, screening, target‑setting, reporting, or research
- Ability to customize fields, thresholds, or taxonomies
- License terms (use in client reporting, benchmarks, products)
- Internal governance: who owns the dataset and validation process
- Total cost vs alternatives, including implementation and maintenance
