Datasets

    Dataset Evaluation Checklist

    A practical guide to help you assess and compare sustainability datasets before integrating them into your investment workflow.

    • Assets, issuers, countries, sectors, time periods covered
    • Point‑in‑time history vs only latest snapshot
    • Percentage of your portfolio/universe covered
    • Treatment of private markets, sovereigns, derivatives
    • Clear documentation of how each metric is calculated
    • Alignment with major frameworks (ISSB, TCFD, SFDR, GHG Protocol, etc.)
    • Handling of estimates vs reported data; use of models and assumptions
    • Version control and change logs for methodology updates
    • Accuracy: Cross‑checks against primary sources or filings
    • Completeness: Missing values, empty fields, coverage by sector/region
    • Consistency: Same definitions over time, across regions and asset classes
    • Outlier handling: Rules for flags, winsorization, or overrides

    Missing data rates

    • % of nulls per field, per sector, per region
    • Are missing values clustered (e.g., EM, small caps, private assets)?

    Coverage vs portfolio

    • % of AUM with non‑missing values for each key metric
    • List of top 20 holdings without data

    Distribution sanity checks

    • Min / max / mean / median for each metric
    • Compare to expected ranges (e.g., CO₂ intensity)
    • Look for impossible values (negative where impossible, >100% where not allowed)

    Outlier detection

    • Z‑scores or percentile bands (e.g., flag values beyond 3σ or 1st/99th percentile)
    • Check whether outliers are data errors or genuine edge cases

    Time‑series stability

    • Year‑on‑year changes; flag jumps above a threshold (e.g., >50% change)
    • Count of restatements or backfills over time

    Cross‑metric consistency

    • Ratios that should reconcile (e.g., Scope 1+2 vs total emissions)
    • Check that derived metrics match their components when you recompute them

    Duplicate & uniqueness checks

    • One record per issuer/ISIN/date where expected
    • No duplicated rows or conflicting values for the same key

    Aggregation checks

    • Aggregate security‑level data and compare with any portfolio‑level numbers provided
    • Recompute vendor‑supplied scores from raw fields on a sample to confirm formulas
    • How often data is refreshed (daily / monthly / annually)
    • Typical lag between real‑world event/report and dataset update
    • Backfill policies when new information becomes available
    • Ability to drill down to source documents or raw inputs
    • Clear audit trail (who changed what, when, and why)
    • Availability of data dictionaries and field‑level metadata
    • Coverage of metrics needed for CSRD/SFDR/Taxonomy, SDR, SEC, ISSB, etc.
    • Mapping tables from dataset fields to reporting templates or PAIs
    • Vendor roadmap for keeping up with new regulations
    • Delivery options: API, SFTP, bulk CSV/Parquet, Excel add‑ins
    • Identifier coverage and mapping (ISIN, LEI, ticker, internal IDs)
    • Performance: file sizes, latency, rate limits, and support for backfills
    • Track record, client base similar to you (AM, pension, insurance)
    • SLAs for uptime, bug fixes, and data corrections
    • Quality of documentation, user guides, and client success support
    • Clear support channels in case of missing files or API disruptions
    • Access to knowledgeable contacts for questions regarding the data, methodology, and underlying assumptions
    • Does the granularity match your needs (security‑level vs scores only)?
    • Suitability for risk, screening, target‑setting, reporting, or research
    • Ability to customize fields, thresholds, or taxonomies
    • License terms (use in client reporting, benchmarks, products)
    • Internal governance: who owns the dataset and validation process
    • Total cost vs alternatives, including implementation and maintenance