Sepehr Dehdashtian
PhD Student in Computer Science
428 S Shaw Ln # 3208
East Lansing, MI
Hi there! I’m Sepehr, a Ph.D. candidate in Computer Science at Michigan State University, advised by Prof. Vishnu N. Boddeti, with an M.Sc. in Electrical Engineering from Sharif University of Technology. My research sits at the intersection of representation learning, generative model evaluation, red-teaming, and controllable generation, with the goal of building AI safety methods for generative and multimodal foundation models: I develop kernel-based representation learning methods that debias and steer model behavior, design evaluations that measure how these models fail, and build automated red-teaming pipelines that expose adversarial vulnerabilities before deployment.
Generative Model Evaluation and Safety Post-training
I’ve done research internships at Apple AIML, Sony AI, and Reality Defender.
- At Apple AIML, I build an automated pipeline that surfaces failure modes in production text-to-image models and their LLM-as-judge evaluators (AutoGraders).
- At Sony AI, I developed post-training methods for safety-oriented fine-tuning and concept erasure in large-scale diffusion models.
- At Reality Defender, I built automated red-teaming pipelines for audio and image deepfake detectors.
Sociotechnical Evaluation and Bias Measurement in Generative Models
I design measurable, distributional definitions and evaluation toolkits for concerns that are usually only described in words, such as stereotypes in text-to-image models (OASIS, ICLR 2025 Spotlight), and I audit how dataset scale affects racial bias in vision-language models (FAccT 2024).
Red-Teaming, Adversarial Robustness, and Controllable Generation
I build automated, black-box red-teaming pipelines that find failure modes in synthetic-image and audio-deepfake detectors (PolyJuice, NeurIPS 2025; FoeGlass, ICML 2026), and training-free steering methods that debias, erase unsafe concepts from, and red-team diffusion models with precise control over the generation process (LatentCompass).
Representation Learning and Kernel Methods for Debiasing
I develop closed-form debiasing methods for vision-language models using kernel methods in reproducing kernel Hilbert spaces (FairerCLIP, ICLR 2024), and I characterize the fundamental trade-offs between utility and fairness in discriminative models (U-FaTE, CVPR 2024).
If you’re interested in discussing my research or anything else, feel free to reach out via email or connect with me on social media. You can find more about my work on this website and in my CV.
I am actively seeking full-time research opportunities and would love to hear from you.
news
| May 11, 2026 | Thrilled to announce that I’ve joined Apple AIML as a Research Intern! 🎉 |
|---|---|
| Apr 30, 2026 | Our paper, FoeGlass: When Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors, has been accepted at ICML 2026. |
| Sep 18, 2025 | Our paper, PolyJuice Makes It Real: Black-Box, Universal Red-Teaming for Synthetic Image Detectors, has been accepted at NeurIPS 2025. |
| Apr 21, 2025 | I’ve been awarded the Interdisciplinary Inquiry and Teaching Fellowship for 2025! 🎉 |
| Feb 11, 2025 | Our paper, OASIS, has been accepted as spotlight ✨ at ICLR 2025. |