Back to projects
Typed Decision Reliability — CFPB Temporal-Shift Feasibility Pilot

Typed Decision Reliability — CFPB Temporal-Shift Feasibility Pilot

A public, reproducible CPU-only research pilot evaluating temporal shift, label quality, duplication and confidence reliability in CFPB debt-collection complaint classification. Includes a TF-IDF/logistic-regression baseline, chronological tests and a documented GO/NO-GO decision; no neural or Laya training is claimed.

Pythonscikit-learnTF-IDFLogistic RegressionJupyter NotebookData AuditingML Evaluation

Project Overview

Typed Decision Reliability is an openly documented machine-learning research feasibility pilot investigating how classification accuracy, probability confidence and selective prediction behave when consumer-complaint data is evaluated across chronological periods. The pilot uses public Consumer Financial Protection Bureau (CFPB) debt-collection complaint narratives and a deliberately lightweight, CPU-only baseline.

Research Question

Can a credible temporal-shift evaluation task be established before investing in more expensive typed-decision or neural-model experiments? The project audits data provenance, label stability, duplication, class imbalance and confidence behavior before deciding whether to proceed.

What Was Implemented

  • Official-source data audit: documented CFPB narrative archive downloads, source manifests, availability checks and original debt-collection issue labels.
  • Chronological evaluation: training on September 2023–June 2024 narratives, validation on July–September 2024, and evaluation across later observed quarters.
  • Duplicate and template analysis: exact-text checks and lexical-family detection to examine cross-period dependence and possible leakage.
  • Reproducible CPU baseline: word unigram/bigram TF-IDF with multinomial logistic regression, compared with simple baselines.
  • Reliability assessment: accuracy, macro-F1, calibration error, Brier score, negative log-likelihood, risk–coverage behavior and confidence-threshold transfer.
  • Research artifacts: reproducibility instructions, plots, metric tables, literature/novelty review, decision log and a documented follow-up experiment protocol.

Selected Pilot Findings

The documented feasibility decision was GO (4 of 4 prespecified conditions), meaning the task was judged suitable for a potential larger experiment. It does not establish that a typed decision model is better, that a confidence gate is deployment-safe, or that the research is novel. The TF-IDF baseline achieved roughly 47.4%–56.0% accuracy across the five primary future quarters, with macro-F1 around 0.349–0.394. Results also highlight changing narrative availability, label imbalance, and the importance of separating repeated templates from novel examples.

Original research plot of classification metrics across chronological CFPB cohortsOriginal temporal evaluation figure from the project's public research repository; not an illustrative or AI-generated image.

Methodology and Limitations

The study uses consumer-submitted or administrative issue labels, not independently adjudicated ground truth. Published narratives are not representative of all complaints, and changing publication rates complicate claims about concept drift. The pilot examined later cohorts; they cannot subsequently be presented as untouched confirmatory holdouts. No Laya or ModernBERT training was performed, and no GPU training or final paper results are claimed.

Tools and Technologies

Python, scikit-learn, TF-IDF, logistic regression, pandas-style tabular analysis, Jupyter Notebook, reproducible research scripts and CFPB public data.

Explore the Work

Source code and research repository · Pilot report · Feasibility decision · Reproduction guide

More projects