Data Sources

Transparency in how we build diaspora intelligence — every dataset, every signal, every method.

40+
Data sources integrated
14
Countries' business registries
99.7%
Data source uptime
How we source data

Built on verified,
public signals.

Sporafind aggregates diaspora intelligence exclusively from public, legally obtainable data sources. We do not purchase, scrape, or infer private data. Every profile in our system is reconstructed from signals that individuals and organisations have voluntarily made public.

Our AI cross-references 40+ distinct data sources — from business registries to academic databases — to build a coherent, de-duplicated, and confidence-scored profile. The result is intelligence that respects privacy while delivering precision.

40+
Active data feeds
14
Daily refresh cycles
3.2B
Data points processed
94%
Average confidence score
Source categories

Every data class we use.

Seven categories of signal intelligence that power every Sporafind profile.

Professional Networks

LinkedIn public profiles, Xing, and professional association directories. We collect signal-level data only — job titles, industries, stated expertise — never private messages or non-public connections.

LinkedIn public data Xing Professional bodies Signal data only
Business Registries

Companies House (UK), SEC EDGAR (US), Ghana Registrar General's Department, CAC (Nigeria), plus 20+ other business registries across Africa, Europe, and North America. Director names, company roles, and filing histories.

Companies House SEC EDGAR Ghana RGD CAC Nigeria 20+ registries
Financial Data

Crunchbase company profiles, PitchBook investor data, SEC filings (13D, 13G, Form 4), Bloomberg terminal feeds, and public market disclosures. Investment history, fund data, and institutional affiliations.

Crunchbase PitchBook SEC filings Bloomberg
Academic Databases

ORCID researcher profiles, Google Scholar publication records, university faculty directories, and ResearchGate. Academic affiliation, field of study, publication history, and institutional roles.

ORCID Google Scholar Faculty directories ResearchGate
Public Records

News media archives (global English, French, Portuguese, and Twi-language outlets), press release distribution networks, and public appointments registers. Media mentions, appointments, and public recognitions.

Media archives Press releases Appointments registers 4-language coverage
Organisational Data

NGO directories (Guidestar, NGO Advisor), diaspora association listings, cultural organisation registries, and community foundation databases. Membership estimates, leadership rosters, and programme portfolios.

Guidestar NGO directories Diaspora associations Cultural registries
Geographic & Demographic Data

Census data from diaspora-hosting nations (UK, US, Canada, Germany, Netherlands, Italy, Spain, South Africa, Australia), diaspora migration statistics from national statistical offices, and World Bank remittance flow data. Geographic distribution, population estimates, migration trends, and remittance corridors.

National census data Migration statistics World Bank remittance data 9+ national statistical offices
AI methodology

How AI processes
raw data into intelligence.

Three layers of machine intelligence transform scattered public signals into structured, actionable diaspora profiles.

Cross-Referencing
Our entity-resolution engine matches records across disparate sources — a LinkedIn profile with a Companies House director record, an ORCID iD with a Crunchbase founder entry. Every connection is verified bidirectionally before it enters the profile graph.
Deduplication
Name variants, transliteration differences (Kwesi / Kweisi), and duplicate listings are resolved through fuzzy matching on multiple identity dimensions — name, email hash, institutional affiliation, geographic location. One person, one profile — always.
Confidence Scoring
Every data point receives a signal-confidence score based on source authority, cross-reference count, recency, and internal consistency. Profiles are ranked by overall confidence — you see 94% average accuracy across the platform, with per-field scores on every profile card.
Privacy by design

What we do NOT collect.

Sporafind is built on a strict data ethics framework. Some categories are permanently out of scope — and we maintain a public audit trail to verify this.

No private messages
We never access, process, or store direct messages, DMs, WhatsApp chats, email inboxes, or any private communication channel.
No non-public financials
We do not access bank account data, private portfolio statements, credit scores, or any non-public financial records.
No biometric data
We do not collect, store, or process fingerprints, facial recognition data, voice prints, or any biometric identifiers.
No browsing history
We do not track, purchase, or infer web browsing history, search history, clickstream data, or device telemetry.

Data freshness & update frequency

We believe transparency includes being precise about data age. Every profile card displays a "last verified" timestamp per source. Our platform-wide update cadence is documented below, and all changes are logged in our public changelog.

Real-time
SEC filings, news feeds
Daily
Business registries, Crunchbase
Weekly
LinkedIn signals, directories
Monthly
Census, NGO databases

Trust the intelligence. Verify the source.

Every Sporafind profile links back to its original sources. Start exploring the diaspora with full transparency.

No credit card · 14-day trial · Cancel anytime