10/19/2026
By Marley O'Neil
The Francis College of Engineering, Department of Electrical and Computer Engineering, invites you to attend a doctoral dissertation defense by Timothy Lawrence Miskell on “Design and Evaluation of Contrastive Learning, Large Language Models and Vision Language Models for Malware Detection.”
Defense Details
- Defense Date: Wednesday, Nov. 4, 2026
- Time: 10 a.m. – noon
- Location: Room 215, Perry Hall, North Campus
Committee Members
- Advisor: Yan Luo, Ph.D., Electrical and Computer Engineering Department, University of Massachusetts Lowell
- Hengyong Yu, Ph.D., Electrical and Computer Engineering Department, University of Massachusetts Lowell
- Orlando Arias, Ph.D., Electrical and Computer Engineering Department, University of Massachusetts Lowell
- Liang-min Wang, Ph.D., Marvell Technology
Abstract
This dissertation investigates the application of self-supervised learning, Large Language Models (LLMs), and Vision Language Models (VLMs) for malware detection using Windows Portable Executable (WinPE) samples. This work addresses the significant challenges associated with limited labeled data, evolving malware behaviors, and the need for models capable of extracting meaningful representations from both static binaries and dynamically generated execution traces.
First, a self-supervised contrastive learning framework is developed for malware classification. By leveraging malware-specific augmentation strategies tailored to the WinPE format, the proposed approach learns robust feature representations from unlabeled samples prior to supervised fine-tuning. Experimental results demonstrate substantial improvements over conventional supervised learning methods in low-label scenarios, achieving macro-averaged F1 scores exceeding 0.86, while significantly reducing dependencies on labeled training data.
Second, this work explores the use of Large Language Models (LLMs) for static malware analysis. A domain-specific representation of Windows Portable Executable (WinPE) files is constructed for parsing and translating headers into a structured textual format suitable for language model processing. Through a series of ablation studies, the proposed DistilBERT-based architecture achieved classification accuracies approaching 100% when trained exclusively on DOS header information, suggesting that highly discriminative malware-family characteristics may be preserved within otherwise overlooked regions of the executable format. These findings indicate that compact structural representations can provide substantial predictive power, while also highlighting the influence of dataset-specific artifacts and software development patterns that warrant further investigation. The results demonstrate that specialized language models can effectively learn malware-relevant semantics from a small fraction of the original binary content.
Finally, a novel VLM-based framework is introduced for dynamic malware analysis. Instead of classifying binaries directly, execution traces are recorded, transformed into video streams, and analyzed using a multi-stage pipeline combining YOLOv8 object detection, Florence-2 scene captioning, and Phi-3 based reasoning. The proposed system extracts visual, spatial, and statistical characteristics from program execution behavior and employs prompt-engineering techniques to improve malware classification. Experimental evaluation demonstrates an F1 score of 0.812, while 3 4 prompt ablation studies achieve F1 scores as high as 0.966. By contrast, baseline LLMs operating directly on execution traces fail to produce meaningful classifications.
Collectively, these contributions demonstrate that modern foundation-model approaches can be successfully adapted to both static and dynamic malware analysis. The results show that self-supervised learning reduces reliance on labeled datasets, domain-adapted LLMs improve semantic understanding of the underlying WinPE structure, and VLMs provide a promising new direction for malware detection based on visual representations of runtime behavior.