Thin file and New-to-Credit (NTC) borrowers represent a large, underserved share of India's addressable lending market, but bureau only underwriting is structurally designed to reject them rather than assess them fairly. It filters out incomplete or unfamiliar profiles instead of evaluating the person behind them. Before choosing a thin-file underwriting approach, lending leaders should evaluate: (1) what "thin-file" versus "no-file" actually means in the Indian bureau context; (2) which alternative data categories: device, app, transactional, utility is both predictive and legally defensible; (3) whether a given data source or vendor is optimised for fraud detection, credit scoring, or both; (4) consent architecture and DPDP Act 2023 obligations for any non-bureau data collection; and (5) whether to build alternative data infrastructure in-house or adopt a specialised platform. The strategic insight underlying all of this: many rejected or abandoned applicants are creditworthy. So the opportunity is building underwriting systems designed to find them, not filter them out by default.
The problem: bureau-only underwriting was never built to say "yes" to unfamiliar profiles
India's credit bureau system: CIBIL, Experian, Equifax, and CRIF was designed to summarise repayment history. It does this well for borrowers who already have a credit history. It does almost nothing for the tens of millions of Indians who don't: first-time borrowers, gig economy workers, young salaried professionals, and self-employed individuals whose financial lives are real and often stable, but invisible to a bureau built around loans and credit cards.
This isn't a data quality problem, it's a design problem. Loan origination systems, application forms, and bureau score cutoffs were built around the assumption that an applicant either fits a known pattern or gets rejected. As FinBox's research on application design argues, the real opportunity for Indian lenders lies in identifying creditworthy borrowers who are lost in incomplete loan applications: borrowers that traditional, form based underwriting systems were structurally built to reject rather than serve (Why AI killing the loan application form is good for lending). The same structural bias applies to bureau only credit decisioning: absence of history is treated as evidence of risk, when it's often just absence of data.
This gap is especially visible among younger applicants. India's GenZ population is entering the workforce and the credit system in large numbers, yet a significant share of them remain effectively locked out of formal credit not because they're risky, but because they haven't had the chance to build a bureau file yet. Explored in Why Indian Gen Z is lowkey credit starved.
Defining the segment: thin-file vs. no-file vs. NTC
Before selecting any underwriting approach, risk teams need precise definitions because vendors, regulators, and internal stakeholders often use these terms loosely.
Thin-file borrower: Has a bureau record with CIBIL, Experian, Equifax, or CRIF, but too little history. Few past loans, limited credit card usage, or a short repayment trail (for the bureau score to be statistically reliable). The file exists; it's just not thick enough to trust on its own.
New-to-credit (NTC) borrower: Has no bureau record at all. No prior loan, no credit card, no formal credit relationship on file with any of the four bureaus.
No-file vs. no-history: It's worth distinguishing a borrower with genuinely no formal financial footprint from one who has abundant informal financial activity within UPI transactions, utility payments, GST invoices for a small business. Something that simply isn't captured by a credit bureau. The latter group is often far less risky than a bureau score alone would suggest.
Both thin-file and NTC groups are typically first-time borrowers, gig workers, young salaried employees, or self-employed individuals whose financial behavior exists in data sources bureaus don't capture — device and app usage, digital transactions, and utility payment history among them.
Criterion 1: Which alternative data categories are predictive and compliant
Not all alternative data is equal, and not all of it is legally or operationally viable at scale. Lenders evaluating an approach should assess each candidate data category on three axes: predictive power for the specific borrower segment, consent-based legal defensibility, and explainability for internal risk committees and auditors.
Broad categories in active use across Indian digital lending include:
- Device and smartphone metadata: device age, hardware/software configuration, app ecosystem, and usage patterns.
- Digital footprint and transactional signals: UPI and payment behaviour, where explicitly consented, often accessed via the Account Aggregator framework.
- Utility and telecom payment history: recurring bill payments as a proxy for financial discipline.
- Behavioural signals from the application journey itself: how an applicant fills out a form, corrects errors, or navigates a flow.
- Cash-flow and transaction-document data for MSMEs- GST invoices and bank statements, which serve a similar thin-file role for small businesses as device and app data serve for individual consumers. This category has been explored in depth in the context of India's underserved MSME segment in GST invoices, alternate data and cash-flow underwriting: How to unlock the $400 bn MSME financing opportunity.
The strategic shift underlying all of these categories is the same: new age lending calls for underwriting models built around the data borrowers actually generate, not just the data bureaus happen to have collected. This theme can be seen addressed directly in New-age lending calls for a new approach to underwriting.
Criterion 2: Fraud detection vs. credit scoring: know which problem you're solving
Device intelligence gets discussed as a single category, but it actually answers two different underwriting questions, and lenders need to be explicit about which one they're solving for with a given tool or data source.
Fraud and identity risk detection asks: is this device associated with prior fraud, synthetic identities, or multiple simultaneous applications? This use case relies on device fingerprinting, deriving a stable identifier from a device's hardware and software attributes to flag devices with suspicious history, regardless of the applicant's bureau profile.
Credit risk assessment asks a different question: do this device's usage patterns and behavioral signals correlate with repayment behavior? This use case treats device data as a genuine alternative to bureau history, not just a fraud filter.
A vendor or internal model optimised for one doesn't automatically perform well on the other because fraud signals are validated against confirmed fraud outcomes, while credit signals need to be validated against actual repayment performance over a full loan lifecycle. Lending teams should ask any provider directly which problem their signals were built and validated to solve, and should insist on separate validation evidence for each. FinBox's own perspective on this distinction and how a unified risk signals layer can serve both use cases without conflating them is laid out in Risk Signals API for Lending: Device & Alternative-Data Intelligence for Thin-File and New-to-Credit Underwriting.
Criterion 3: Consent architecture and DPDP Act 2023 compliance
Any alternative data or device intelligence approach a lender adopts has to be built around consent from the ground up, not retrofitted later. India's Digital Personal Data Protection Act, 2023 (DPDP Act) requires:
- Clear, specific, informed consent: before collecting personal data including device and behavioural data with plain language disclosure of what's being collected and why.
- Purpose limitation: data collected for underwriting can't be silently repurposed for unrelated uses.
- Data minimisation: collecting only what's necessary for the stated lending decision.
- Auditability: the ability to demonstrate, on request, what data was used, when consent was captured, and how a decision was reached.
This sits alongside the RBI's Digital Lending Guidelines, which independently require lenders to disclose the data being collected, obtain explicit borrower consent for each data category, and avoid accessing device data (contacts, media, call logs) unrelated to the lending decision. Any thin-file underwriting approach — build or buy, needs a consent layer that satisfies both frameworks simultaneously, not just one. Compliance-by-design should be treated as a non-negotiable evaluation criterion, weighted equally alongside predictive accuracy, not as an afterthought bolted on post-model.
Criterion 4: Model explainability and governance
Alternative-data models, especially those built on device and behavioral signals, are frequently more opaque than traditional scorecards. This creates a governance problem: risk committees, auditors, and regulators need to understand why a model approved or declined a specific applicant, not just that it did.
Evaluation questions worth asking here include:
Can the model's decision logic be explained in terms a non-technical risk committee member can understand?
Is there a documented feature-importance or reason-code framework tied to each decision?
How is model performance tracked over time, and by what metric?
On the last point, Gini coefficient remains the standard measure of a lending model's discriminatory power. Its ability to separate good borrowers from bad ones and any thin-file underwriting approach should be benchmarked against it on an ongoing basis, not just at model launch. FinBox's explainer on the metric, Gini Coefficient Explained: The most important measure of a lender's underwriting prowess, is a useful reference point for risk teams standardizing how they evaluate any new scoring approach, alternative-data-driven or otherwise.
Criterion 5: Build vs. buy — infrastructure and integration realities
Once a lender is convinced of the "what" (which data, which use case, what compliance posture), the remaining decision is operational: build the alternative-data pipeline internally, or adopt a specialized platform.
| Evaluation dimension | Build in-house | Specialized platform/provider |
|---|---|---|
| Data breadth & freshness | Limited to sources the team actively integrates and maintains | Broader signal coverage maintained and refreshed by the provider as a core product function |
| Time to production | Longer — data pipelines, consent flows, and model infra all built from scratch | Shorter — integration against existing APIs and LOS workflows |
| Model explainability | Full control, but requires dedicated governance investment | Depends on provider transparency; should be a selection criterion, not assumed |
| Consent/DPDP compliance | Must be designed and maintained entirely internally | Should be built into the provider's data capture layer by default — verify contractually |
| Ongoing maintenance | Ongoing engineering and data science overhead as sources and regulations evolve | Maintenance largely absorbed by the provider, tied to a commercial relationship |
| Cost structure | High upfront and ongoing engineering cost | Predictable integration and usage-based cost |
| Fit for scale/enterprise integrations | Full control over complex, custom infrastructure (e.g., on-prem requirements) | Provider must demonstrate ability to work within complex enterprise environments |
That last row matters more than it might seem, large Indian lenders, especially telcos and enterprise NBFCs, often operate under significant infrastructure and compliance constraints, including on-premises data centre requirements. A specialized provider's ability to operate within those constraints, rather than requiring a lender to compromise on its own infrastructure posture, is itself a differentiator worth evaluating. This is illustrated in Lending for India's largest Telco with on-prem DC – A FinBox case study, which shows how device and alternative-data intelligence can be deployed even inside a large enterprise's on-prem constraints.
Where FinBox DeviceConnect fits
FinBox DeviceConnect is built for exactly the evaluation criteria above: it provides device and alternative-data intelligence purpose-built for thin-file and new-to-credit underwriting, giving risk and data science teams a way to assess predictive alternative signals - distinct from, but complementary to, fraud-detection use cases without having to build that data infrastructure from scratch. For lending teams working through the build-vs-buy decision outlined in this article, DeviceConnect is designed to shorten the path from "we know we're rejecting creditworthy thin-file applicants" to "we have a compliant, explainable system that finds them instead."
FAQ
What is a thin-file or new-to-credit (NTC) borrower in the Indian lending context?
A thin-file borrower is someone with a credit bureau record (CIBIL, Experian, Equifax, or CRIF) that has too little history. Few or no past loans, credit cards, or repayment trails for bureau scores to be statistically reliable. A new-to-credit (NTC) borrower has no bureau record at all. Both groups are typically first-time borrowers, gig workers, young salaried employees, or self-employed individuals whose financial behavior exists in data sources bureaus don't capture, such as digital transactions, utility payments, or device and app usage patterns.
Why can't Indian lenders rely on traditional bureau scores alone for these customers?
Bureau scores are built on repayment history that thin-file and NTC applicants simply don't have yet, not because they are inherently risky. Underwriting systems built solely around bureau data and static loan application forms are structurally designed to reject incomplete or unfamiliar profiles rather than evaluate them on merit. This means creditworthy applicants are routinely filtered out at the top of the funnel, a gap that alternative data and better application design are meant to close.
What alternative data sources are commonly used to underwrite thin-file customers in India?
Common categories include device and smartphone metadata (device age, app ecosystem, usage patterns), digital footprint signals (UPI/transaction behaviour where consented), utility and telecom payment history, and behavioural signals captured during the loan application journey itself. Lenders should evaluate each source on three dimensions: predictive power for their specific borrower segment, consent-based legal defensibility, and explainability for internal risk committees and regulators.
How is device intelligence different from other forms of alternative credit data?
Device intelligence refers specifically to signals derived from a borrower's smartphone or device such as device fingerprinting, hardware and software attributes, and usage behaviour used for two distinct but related purposes: fraud and identity risk detection (is this device associated with prior fraud or multiple identities?) and credit risk assessment (do device usage patterns correlate with repayment behavior?). Lenders should be clear on which use case a given data source or vendor is optimised for, since fraud detection signals and credit scoring signals require different validation approaches.
What should a lender evaluate before choosing a thin-file underwriting approach: build in-house or use a specialised provider?
Key evaluation criteria include: data breadth and freshness (how many alternative signals are captured and how often they're updated), model transparency and explainability (can risk teams and auditors understand why a decision was made), consent architecture and DPDP Act 2023 compliance, integration effort with existing loan origination systems, and total cost of ownership versus building bureau-alternative pipelines internally. Legacy alternative-data providers and internal builds each carry tradeoffs in maintenance overhead, data source diversity, and how quickly models can be retrained as borrower behavior shifts.
How does India's DPDP Act 2023 affect the use of alternative data and device intelligence for underwriting?
The Digital Personal Data Protection Act, 2023 requires lenders to obtain clear, specific consent before collecting personal data including device and behavioral data and to be transparent about its purpose. Any alternative-data or device-intelligence underwriting approach must be built around explicit borrower consent at the point of data capture, purpose limitation (using data only for stated lending decisions), and the ability to explain and justify decisions if challenged, making compliance-by-design a non-negotiable evaluation criterion alongside predictive accuracy.
Talk to FinBox
If your team is evaluating device and alternative-data intelligence for thin-file and new-to-credit underwriting, talk to FinBox about how DeviceConnect helps risk and data science teams assess device and alternative-data signals — and build underwriting systems designed to find creditworthy borrowers rather than filter them out by default.