Few decisions in retail banking have a greater impact than credit scoring. Each underwriting decision affects portfolio performance, customer access to credit, and regulatory scrutiny. Yet a weak credit score does not necessarily signal an unworthy borrower. Many consumers with thin credit files or unimpressive scores may well be financially responsible but just disadvantaged by limited borrowing history, irregular income patterns, or temporary financial setbacks.
For such consumers, a declined loan can have far-reaching implications. They may have to postpone buying a home, alter their education plan, rethink a business, or defer an unexpected yet important expense. Alternative data can help avoid such a scenario by providing additional information for a more complete view of an applicant’s financial behaviour and ability to repay.
We discuss how consented alternative data can improve risk differentiation for thin-file (or unscored) applicants. But for this approach to work as intended, financial institutions must focus on establishing strong governance and controls around consent, data minimisation, security, fairness testing, validation, and monitoring.
Before moving any further, it is worth examining why thin-file customers continue to present a challenge for traditional credit-scoring approaches. Credit decisions historically relied on relationship-based judgment and limited documentation. The industry moved from largely judgment-based lending to data-driven underwriting decades ago. What has changed recently is not the objective of credit scoring, but the volume of customer data available to support decisions for applicants who have little or no traditional credit history.
When evaluating the potential of alternative data, most risk teams focus on a few practical questions: Are approval rates improving? Are the losses within expectations? Can the decision be defended?
Interestingly, even as new data sources emerge, bureau scores remain the foundation of most credit decisions, especially in the retail segment, because they are well-understood and supported by decades of performance history. The gap is the thin-file (or unscored) population—customers who may be creditworthy but lack sufficient bureau history to be scored or be priced confidently.
According to Consumer Financial Protection Bureau’s Technical correction and update to the CFPB’s credit invisibles estimate, June 2025, approximately 7 million people in the US were credit invisible in 2020, representing 2.7% of the US adult population. Banks typically use bureau data (e.g., Experian, Equifax, TransUnion in the US; CIBIL in India) as the main underwriting input. Where this data is limited, alternative data can be used, but only if it is lawful, explainable, and independently validated and audited.
Fintech lenders have reported positive results from incorporating alternative data into underwriting decisions. However, banks operate under a stiff regulatory and governance environment, which means their use of alternative data in credit decisions requires a much more through approach.
It is worth examining here the type of alternative data that would prove useful to lenders. A ’social graph’ (derived from an applicant’s network or community interactions), for instance, may be mathematically elegant but not very useful in making lending decisions because it can behave like a proxy for sensitive attributes and raise privacy concerns. The safer option for traditional financial institutions is to focus on direct, consented, repayment-relevant signals (e.g., verified cashflows and payment behaviour) with a clear data lineage.
When opting for alternative data, banks must:
Before alternative data can be adopted, traditional financial institutions need to be confident that customer consent is duly secured and properly documented, data collection remains proportionate to business needs, security controls are effective, and fairness risks have been appropriately assessed. Independent validation and ongoing monitoring are equally important because model behaviour can evolve over time.
Let us look at some potential risks that banks must cover for:
Fair-lending or discrimination risk: One of the biggest concerns with alternative data is that seemingly harmless variables can unintentionally mirror protected characteristics, creating fair-lending risks that are difficult to detect without rigorous testing.
Privacy and security risk: The more sensitive the data, the higher the impact of a breach and regulatory exposure. This makes controls such as explicit consent, data minimisation, encryption and access controls, vendor due diligence, retention limits, and incident-response readiness extremely crucial considerations.
Model integrity risk (gaming or drift): Some digital signals are easier to manipulate and may decay quickly. To counter this, financial institutions should lean on verifiable sources, conduct robustness testing, use challenger models, monitor drift, and recalibrate the models periodically.
For decades, credit scoring has helped lenders balance growth and risk, yet millions of potential customers remain beyond its reach. Lenders often face the challenge of expanding approval rates for applicants with limited credit history – the thin-file applicants. While alternative data can play an important role, its success depends more on governance than on the data itself.
To go mainstream with alternative data, financial institutions must focus on a three-step approach:
Finally, there must be a clear mechanism around monitoring and cadence; early warning indicators merit close attention during the first few months of any pilot. Data-quality issues, unexpected customer behaviour, or shifts in model performance can emerge quickly and are far easier to address before the solution is deployed at scale.
Before closing, here’s a quick word on some common pitfalls that financial institutions must be aware of when integrating alternative data into their credit scoring models. Using weak proxies (e.g., social popularity) introduces a privacy or fairness risk, and must be avoided at all costs unless they offer a strong and credible predictive value. It is also important to refrain from expanding the scope before controls (such as consent capture, monitoring, validation, and escalation paths) are proven.
Banks should avoid relying on vendor-reported uplift without institution-specific back-testing and pilot measurement, and keep the production models free of data with opaque lineage and/or uncontrolled data changes (schema shifts, missing feeds). Lastly, it would be unwise to treat fairness as a one-time exercise – continuous monitoring is key.