10/02/2026
By Namrata Shivagunde

The Kennedy College of Sciences' Miner School of Computer and Information Sciences invites you to attend a doctoral defense by Namrata Shivagunde titled "Low-Rank Pre-Training and Beyond: Multi-Perspective Diagnostics for Structural and Behavioral Equivalence in Large Language Models."

  • Date: Oct. 16, 2026
  • Time: 1:30 – 3 p.m.
  • Location:Zoom meeting

Committee members:

  • Anna Rumshisky (advisor), associate professor, Miner School of Computer & Information Sciences
  • Benyuan Liu, professor, Miner School of Computer & Information Sciences
  • Ming Shao, associate professor, Miner School of Computer & Information Sciences
  • Giannis Karamanolakis, senior applied scientist, Amazon AGI

Abstract

Large language model (LLM) research frequently relies on scalar aggregate metrics such as validation loss, perplexity, and accuracy as proxies for model quality. However, these aggregate measures reduce multi-dimensional LLM behaviors into single values, creating an illusion of equivalence between models that optimize, represent, and generalize in fundamentally different ways. This metric-centric view is further compromised by inconsistent experimental setups, noisy small-scale probes, and limited prompt sensitivity assessments.

In this thesis, we address these evaluation gaps across three core directions: benchmarking and diagnosing the internal mechanics of low-rank pre-training, scaling psycholinguistic probes to assess linguistic competence reliably, and deconstructing in-context prompt sensitivity under controlled corruptions.

First, to establish empirical parity in low-rank pre-training methods, we systematically benchmark low-rank model decomposition methods, such as ReLoRA, and memory-efficient optimizers, such as GaLore, under a strictly identical hardware setup. Within this unified testbed, we introduce three architectural improvements to ReLoRA-style methods and two to GaLore-style methods, achieving a smaller memory footprint at full-rank throughput parity, with validation perplexity close to full-rank training.

We then show that matching validation perplexity does not imply structural equivalence. We introduce a comprehensive diagnostic framework that evaluates five low-rank pre-training methods across four complementary dimensions using 16 metrics: 1-D loss landscape curvature along random and top-1 PCA directions, trajectory interpolation between checkpoints, spectral structure of weights and updates, and activation similarity. We demonstrate that each method settles into a geometrically distinct basin, activation divergence increases in later layers, and incorporating geometric and spectral metrics improves downstream performance prediction over perplexity alone.

Extending this diagnostic paradigm to linguistic capabilities, we address the limitations of small-scale psycholinguistic probes by constructing scaled, statistically reliable benchmarks for negation (NEG-1500) and role reversal (ROLE-1500). Evaluating 22 models reveals that performance drops 20–57% relative to legacy small datasets, demonstrating that small-sample evaluations skew behavioral claims and mask underlying model insensitivity to core linguistic structures.

To investigate inference-time behavior, we systematically deconstruct in-context learning by corrupting four prompt components — task descriptions, demonstration inputs, labels, and inline instructions — across models ranging from 1.5B to 70B parameters across 10 diverse datasets. We find that repeating text within the prompt boosts model performance, and larger models (greater than 30B) are more sensitive to the semantics of the prompt.