About the Role
This is a founding research role at the center of a new effort to reimagine consumer credit scoring from the model architecture up. You will own the foundation models that underpin the entire research program, designing and pre-training purpose-built encoder architectures for tabular and event-sequence credit data. The work is high-stakes, technically deep, and shapes the direction of the team from day one.
What You'll Do
-
Design, implement, and pre-train encoder-first transformer architectures for heterogeneous tabular and sequential credit-file data, spanning tokenization schemes through training objectives.
-
Drive the research agenda on representation learning for credit, including self-supervised objectives, handling of missingness and censoring, temporal drift, and transfer to downstream tasks.
-
Own the training infrastructure end-to-end: distributed training, experiment tracking, evaluation harnesses, and rigorous ablation discipline.
-
Collaborate with causal and explainability researchers to build interpretability into architectures by construction, through attention structures, bottlenecks, and monotonicity constraints.
-
Benchmark rigorously against gradient-boosting incumbents and published tabular foundation models, and communicate results to internal, regulatory, and external audiences.
What We're Looking For
-
8 to 10 or more years of experience, with at least 5 years specifically building and training transformer encoder architectures from scratch, including attention mechanisms, positional and temporal encodings, and tokenization of non-text data at scale.
-
Demonstrated experience pre-training foundation models on large datasets (millions or more of examples) using distributed training infrastructure such as PyTorch Distributed, DeepSpeed, or equivalent.
-
Deep familiarity with tabular and sequence foundation model literature and a well-formed, defensible view of its limitations.
-
Strong hands-on proficiency in Python and PyTorch, including CUDA-level distributed training and experiment tracking systems.
-
Experience designing evaluation harnesses, ablation studies, and reproducible benchmarking pipelines that move models from research notebooks to multi-GPU training.
-
Background with event-sequence or time-series transformers on transaction, clinical, or clickstream data is a strong plus.
-
Prior work in credit, lending, financial services, or other regulated domains where model governance and explainability constrained architecture choices is highly valued.
-
Publications at top-tier ML venues such as NeurIPS, ICML, ICLR, or KDD, or equivalent significant open-source contributions to foundation model research.
Location
This is an on-site role. Primary location is San Francisco, CA, with additional office options in New York, NY and Washington, DC.