Aniket Roy

Aniket Roy

Researcher, Media Analytics Department

NEC Laboratories America

About

I am a Researcher in the Media Analytics Department at NEC Laboratories America, where my work focuses on computer vision, multimodal learning, and generative models for large-scale visual understanding.

My research investigates how machines can interpret and reason about images and video, with particular emphasis on vision-language learning, data-efficient recognition, and generative techniques that enhance visual perception and representation. My broader goal is to advance next-generation visual intelligence systems for media analytics and multimodal machine learning.

I received my Ph.D. in Computer Science from Johns Hopkins University, advised by Prof. Rama Chellappa, and my M.S. (by Research) in Computer Science and Engineering from the Indian Institute of Technology Kharagpur. My work has appeared at leading venues including CVPR, NeurIPS, ICCV, and WACV. I was invited to the ICCV 2025 and AAAI 2026 Doctoral Consortiums and was named an Amazon Fellow through the JHU + Amazon Initiative for Interactive AI. During my doctoral studies I held research internships at Qualcomm, SRI International, Amazon AWS AI, and MERL.

Vision-Language Learning Generative Models Few-Shot Learning Multimodal Representation Learning Multimedia Security

News

Publications

* denotes equal contribution.

Generative Models & Personalization
KC-3DGS
KC-3DGS: Kurtosis-Constrained Gaussian Splatting for High-Fidelity View Synthesis
Vivekjyoti Banerjee, Abhay Yadav, Rama Chellappa, Aniket Roy
Preprint 2026

Wavelet-domain supervision for 3D Gaussian Splatting, combining multi-scale coefficient alignment, kurtosis concentration, and cross-band covariance penalties to improve rendering quality, especially from sparse views.

AAAI-26 Doctoral Consortium poster
Learning More from Less: Resource-Constrained Generative AI for Classification, Generation, and Personalization
Aniket Roy
AAAI 2026Doctoral Consortium

An overview of my doctoral research on resource-efficient generative AI: uncertainty-guided mixup for few-shot learning, diffusion models grounded in natural image statistics, and parameter-efficient low-rank adapters for personalized synthesis.

DuoLoRA
DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style Personalization
Aniket Roy, Shubhankar Borse, Shreya Kadambi, Risheek Garrepalli, Debasmit Das, Shweta Mahajan, Hyojin Park, Ankita Nayak, Rama Chellappa, Munawar Hayat, Fatih Porikli
ICCV 2025

Content-style personalization of diffusion models via cycle-consistent training and layer-wise diffusion priors.

MultLFG
MultLFG: Training-Free Multi-LoRA Composition Using Frequency-Domain Guidance
Aniket Roy, Maitreya Suin, Ketul Shah, Rama Chellappa
Preprint

A training-free method for composing multiple LoRA adapters using frequency-domain guidance during diffusion sampling.

DiffNat
DiffNat: Improving Diffusion Image Quality Using Natural Image Statistics
Aniket Roy, Maitreya Suin, Anshul Shah, Ketul Shah, Jiang Liu, Rama Chellappa
TMLR 2025

A kurtosis-based loss grounded in natural image statistics that improves the perceptual quality of diffusion-generated images.

Diffuse2Adapt
Diffuse2Adapt: Controlled Diffusion for Synthetic-to-Real Domain Adaptation
Ketul Shah, Arushi Sinha, Arun Reddy, Aniket Roy, Rama Chellappa
ICIP 2025

Controlled diffusion conditioned on target-domain context and style turns synthetic renders into realistic training images, raising VisDA target accuracy from 90.3% to 91.8%.

DiversiNet
DiversiNet: Mitigating Bias in Deep Classification Networks across Sensitive Attributes through Diffusion-Generated Data
Basudha Pal, Aniket Roy, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa
IJCB 2024

Uses diffusion-generated data to reduce demographic bias in deep classifiers.

Few-Shot, Multimodal & Representation Learning
AeroGen
AeroGen: Ground-to-Air Generalization for Action Recognition
Ketul Shah, Anshul Shah, Arun Reddy, Aniket Roy, Celso M. de Melo, Rama Chellappa
FG 2025

Generalizing action recognition from ground-level video to aerial viewpoints.

Cap2Aug
Cap2Aug: Caption-Guided Image Data Augmentation
Aniket Roy, Anshul Shah*, Ketul Shah*, Anirban Roy, Rama Chellappa
WACV 2025MAR Workshop @ CVPR 2025

Semantic data augmentation for low-data regimes using pretrained captioning and text-to-image diffusion models.

HaLP
HaLP: Hallucinating Latent Positives for Skeleton-Based Self-Supervised Learning of Actions
Anshul Shah, Aniket Roy*, Ketul Shah*, Shlok Mishra, David Jacobs, Anoop Cherian, Rama Chellappa
CVPR 2023

Generates hard latent positives to improve contrastive self-supervised learning of skeleton-based action encoders.

Certified robustness
Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz Regularization
Mahyar Fazlyab*, Taha Entesari*, Aniket Roy, Rama Chellappa
NeurIPS 2023

A differentiable regularizer that lower-bounds the distance of data points to the decision boundary, improving certified robustness.

FeLMi
FeLMi: Few-Shot Learning with Hard Mixup
Aniket Roy, Anshul Shah, Ketul Shah, Prithviraj Dhar, Anoop Cherian, Rama Chellappa
NeurIPS 2022

Improves few-shot classification by generating hard mixup samples that sharpen decision boundaries.

DiffAlign
DiffAlign: Few-Shot Learning Using Diffusion-Based Synthesis and Alignment
Aniket Roy, Anshul Shah*, Ketul Shah*, Anirban Roy, Rama Chellappa
Preprint

Leverages text-to-image diffusion models to synthesize and align training data for few-shot recognition.

MuLOT
Multimodal Learning Using Optimal Transport for Sarcasm and Humor Detection
Shraman Pramanick*, Aniket Roy*, Vishal M. Patel
WACV 2022

Optimal-transport-based fusion of video, audio, and text for detecting sarcasm and humor.

BRI3L
BRI3L: A Brightness Illusion Image Dataset for Identification and Localization of Regions of Illusory Perception
Aniket Roy, Anirban Roy, Soma Mitra, Kuntal Ghosh
ICIP 2024

A benchmark dataset for brightness illusions, with an analysis of how diffusion models perceive illusory regions.

Fairness in Face Recognition
PASS
PASS: Protected Attribute Suppression System for Mitigating Bias in Face Recognition
Prithviraj Dhar*, Joshua Gleason*, Aniket Roy, Carlos D. Castillo, Rama Chellappa
ICCV 2021

An adversarial framework that suppresses protected attributes in face descriptors to reduce demographic bias.

Distill and De-bias
Distill and De-bias: Mitigating Bias in Face Recognition Using Knowledge Distillation
Prithviraj Dhar, Joshua Gleason, Aniket Roy, Carlos D. Castillo, P. Jonathon Phillips, Rama Chellappa
arXiv

A knowledge-distillation framework for reducing demographic bias in face recognition.

Multimedia Security & Forensics
Digital Image Forensics book
Digital Image Forensics: Theory and Implementation
Aniket Roy, Rahul Dixit, Ruchira Naskar, Rajat Subhra Chakraborty
BookSpringer, ISBN 978-9811076435

A monograph covering the theory and practice of digital image forensics.

Optimal PEE watermarking
Towards Optimal Prediction Error Expansion Based Reversible Image Watermarking
Aniket Roy, Rajat Subhra Chakraborty
IEEE TCSVT

A theoretical analysis of prediction-error-expansion reversible watermarking toward optimal embedding.

CG vs natural image classification
Classification of Computer-Generated and Natural Images Based on Efficient Deep Convolutional Recurrent Attention Model
Diangarti Tariang, Prithviraj Sengupta, Aniket Roy, Rajat Subhra Chakraborty, Ruchira Naskar
CVPRW 2019

A recurrent attention model for distinguishing computer-generated from natural images.

DCTR forgery detection
Discrete Cosine Transform Residual Feature Based Filtering Forgery and Splicing Detection in JPEG Images
Aniket Roy, Diangarti Tariang, Rajat Subhra Chakraborty, Ruchira Naskar
CVPRW 2018

DCT residual features for detecting filtering forgery and splicing in JPEG images.

Camera source identification
Camera Source Identification Using Discrete Cosine Transform Residue Features and Ensemble Classifier
Aniket Roy, Rajat Subhra Chakraborty, Udaya Sameer, Ruchira Naskar
CVPRW 2017

Camera source identification using DCT residue features with an ensemble classifier.

Copy-move forgery detection
Copy-Move Forgery Detection with Similar but Genuine Objects
Aniket Roy, Akhil Konda, Rajat Subhra Chakraborty
ICIP 2017

Rotated local binary pattern features to separate copy-move forgeries from visually similar genuine objects.

JPEG forgery detection
Automated JPEG Forgery Detection with Correlation-Based Localization
Diangarti Tariang, Aniket Roy, Rajat Subhra Chakraborty, Ruchira Naskar
ICMEW 2017

Detection and localization of JPEG forgeries based on noise correlation.

Optimal distortion estimation
Optimal Distortion Estimation for Prediction Error Expansion Based Reversible Watermarking
Aniket Roy, Rajat Subhra Chakraborty
IWDW 2016Best Paper Award

Theoretical analysis of embedding distortion in prediction-error-expansion reversible watermarking.

HVS watermarking
An HVS-Inspired Robust Non-Blind Watermarking Scheme in YCbCr Space Designed to Prevent Singular Value Exchange Attacks
Aniket Roy, Arpan Kumar Maiti, Kuntal Ghosh
IJIGInternational Journal of Image and Graphics

A human-visual-system-inspired watermarking scheme with a robustness analysis.

Reversible color image watermarking
Reversible Color Image Watermarking in the YCoCg-R Color Space
Aniket Roy, Rajat Subhra Chakraborty, Ruchira Naskar
ICISS 2015

A reversible color image watermarking scheme with an information-theoretic analysis.

Honors & Awards