Publications
2026
Internal Safety Collapse in Frontier Large Language Models [Code]
. NeurIPS, Sydney, Australia, 2026.
VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation [Code]
. NeurIPS, Sydney, Australia, 2026.
ShadowFPT: Backdooring Federated Prompt Tuning via Shadow Triggers
. NeurIPS, Sydney, Australia, 2026.
Towards Multi-Human-Value Alignment via Value Localization in LLMs
. NeurIPS, Sydney, Australia, 2026.
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents [Code]
. MM, Rio de Janeiro, Brazil, 2026.
RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion
. MM, Rio de Janeiro, Brazil, 2026.
BackdoorVLM: A Benchmark for Backdoor Attacks and Defenses on Vision-Language Models [Code]
. MM, Rio de Janeiro, Brazil, 2026.
Agent4POI: Agentic context-conditioned affordance reasoning for Multimodal Point-of-Interest Recommendation
. MM, Rio de Janeiro, Brazil, 2026.
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
. IROS, Pittsburgh, Pennsylvania, USA, 2026.
Perturbation Effects on Robustness and Individual Fairness
. KDD, Jeju Island, South Korea, 2026.
FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and Content
. ICML, Seoul, South Korea, 2026.
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs [Code] [Project]
. ICML, Seoul, South Korea, 2026.
SciAgentGym: Benchmarking Multi-Step Scientific Tool-Use in LLM Agents
. ICML, Seoul, South Korea, 2026.
AudioMosaic: Contrastive Masked Audio Representation Learning
. ICML, Seoul, South Korea, 2026.
Towards Context-Invariant Safety Alignment for Large Language Models
. ICML, Seoul, South Korea, 2026.
MESA: Improving MoE Safety Alignment via Decentralized Expertise
. ICML, Seoul, South Korea, 2026.
RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry [Code]
. ICML, Seoul, South Korea, 2026.
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints
. ACL, San Diego, California, USA, 2026. [Main, Oral]
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents [Code]
. ACL, San Diego, California, USA, 2026. [Findings]
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
. ACL, San Diego, California, USA, 2026. [Findings]
OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens [Code] [Project Page] [Hugging Face]
. CVPR, Denver CO, USA, 2026.
GenBreak: Red Teaming Text-to-Image Generation Using Large Language Models
. CVPR, Denver CO, USA, 2026.
WithAnyone: Towards controllable and id consistent image generation [Code] [Project]
. ICLR, Rio de Janeiro, Brazil, 2026.
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models [Code]
. ICLR, Rio de Janeiro, Brazil, 2026.
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
. AAAI, Singapore, 2026.
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
. WWW, Dubai, United Arab Emirates, 2026.
Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
. AAAI, Singapore, 2026.
Do We Really Need SFT? Prompt-as-Policy over Knowledge Graphs for Cold-start Next POI Recommendation
. CIKM, Rome, Italy, 2026.
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026.
FedEGG: Federated Learning with Explicit Global Guidance
. Frontiers of Computer Science (FCS), 2026.
On the Adversarial Transferability of Generalized "Skip Connections" [Code]
. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026.
OpenRedRL: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming [Code]
. Frontiers of Computer Science (FCS), 2026.
Defense-to-attack: Bypassing weak defenses enables stronger jailbreaks in Vision-Language Models
. Pattern Recognition, 2026.
Learnable Coreset Selection for Graph Active Learning
. Transactions on Machine Learning Research (TMLR), 2026.
2025
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models [Code]
. NeurIPS, San Diego, USA, 2025.
SafeVid: Toward Safety Aligned Video Large Multimodal Models [Dataset]
. NeurIPS, San Diego, USA, 2025.
OmniSVG: A Unified Scalable Vector Graphics Generation Model [Code] [Project Page] [Hugging Face]
. NeurIPS, San Diego, USA, 2025.
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
. NeurIPS, San Diego, USA, 2025.
JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
. NeurIPS, San Diego, USA, 2025.
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety[Code]
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks[Code]
. TDSC, 2025.
BadPatch: Diffusion-Based Generation of Physical Adversarial Patches[Code] [AdvT-shirt-1K Dataset]
. ICCV Workshop Findings, Honolulu, Hawai'i, 2025.
IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves[Code]
. ICCV, Honolulu, Hawai'i, 2025.
Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation
. ICCV, Honolulu, Hawai'i, 2025.
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
. ICCV, Honolulu, Hawai'i, 2025.
T2UE: Generating Unlearnable Examples from Text Descriptions
. MM, Dublin, Ireland, 2025.
FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models
. MM, Dublin, Ireland, 2025.
From Failures to Fixes: LLM-Driven Scenario Repair for Self-Evolving Autonomous Driving[Code]
. MM, Dublin, Ireland, 2025.
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos[Code]
. Dataset Track, MM, Dublin, Ireland, 2025.
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP[Code] [Project] [HuggingFace demo]
. ICML, Vancouver, Canada, 2025.
Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks[Project]
. CVPR, Nashville TN, 2025.
TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models[Code]
. CVPR, Nashville TN, 2025.
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-Language Models[Code] [Project]
. CVPR, Nashville TN, 2025.
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks[Code]
. ICLR, Singapore, 2025.
Detecting Backdoor Samples in Contrastive Language Image Pretraining[Code] [Project]
. ICLR, Singapore, 2025.
AIM: Additional Image Guided Generation of Transferable Adversarial Attacks
. AAAI, Philadelphia, USA, 2025.
CALM: Curiosity-Driven Auditing for Large Language Models[Code]
. AAAI, Philadelphia, USA, 2025.
HoneypotNet: Backdoor Attacks Against Model Extraction
. AAAI, Philadelphia, USA, 2025.
MMFair: Fair Learning via Min-Min Optimization
. CIKM, Seoul, Korea, 2025.
Optimizing Cross-Client Domain Coverage for Federated Instruction Tuning of Large Language Models
. EMNLP 2025 Findings, Suzhou, China, 2025.
Learning from Heterogeneity: A Dynamic Learning Framework for Hypergraphs [Code]
. TAI, To appear in 2025.
2024
UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
. NeurIPS, Vancouver, Canada, 2024.
ModelLock: Locking Your Model With a Spell
. MM, Melbourne, Australia, 2024.
White-box Multimodal Jailbreaks Against Large Vision-Language Models
. MM, Melbourne, Australia, 2024.
AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning [Code]
. MM, Melbourne, Australia, 2024.
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models [Code]
. MM, Melbourne, Australia, 2024.
Adversarial Prompt Tuning for Vision-Language Models [Code]
. ECCV, MiCo Milano, Italy, 2024.
Constrained Intrinsic Motivation for Reinforcement Learning [Code]
. IJCAI, Jeju, Korea, 2024.
Toward Evaluating Robustness of Reinforcement Learning with Adversarial Policy [Code]
. DSN, Brisbane, Australia, 2024.
VeriFi: Towards Verifiable Federated Unlearning
. TDSC, 2024. (Best Paper Runner-up)
Fake Alignment: Are LLMs Really Aligned Well?
. NAACL, Mexico City, Mexico, 2024.
LDReg: Local Dimensionality Regularized Self-Supervised Learning [Code]
. ICLR, Vienna, Austria, 2024.
Unlearnable Examples For Time Series
. PAKDD, 2024.
2023
Reconstructive Neuron Pruning for Backdoor Defense [Code]
. ICML, Hawaii, USA, 2023.
Unlearnable Clusters: Towards Label-agnostic Unlearnable Examples [Code]
. CVPR, Vancouver, Canada, 2023.
Distilling Cognitive Backdoor Patterns within an Image [Code]
. ICLR, Kigali, Rwanda, 2023.
Transferable Unlearnable Examples[Code]
. ICLR, Kigali, Rwanda, 2023.
On the Importance of Spatial Relations for Few-shot Action Recognition
. MM, Ottawa, Canada, 2023.
Backdoor Attacks on Time Series: A Generative Approach[Code]
. SaTML, 2023.
Relationships between tail entropies and local intrinsic dimensionality and their use for estimation and feature representation
. Information Systems (2023): 102245.
Imbalanced Gradients: A Subtle Cause of Overestimated Adversarial Robustness[Code]
. Machine Learning (2023): 1-26.
Query-efficient Black-box Adversarial Attacks on Automatic Speech Recognition
. To appear in TASLP.
2022
Local Intrinsic Dimensionality, Entropy and Statistical Divergences
. Entropy 24(9), 1220, 2022.
Few-Shot Backdoor Attacks on Visual Object Tracking [Code]
. ICLR, 2022.
CalFAT: Calibrated Federated Adversarial Training with Label Skewness
. NeurIPS, 2022.
Copy, Right? A Testing Framework for Copyright Protection of Deep Learning Models [Code]
. Oakland, 2022.
Backdoor Attacks on Crowd Counting [Code]
. MM, 2022.
Fine-mixing: Mitigating Backdoors in Fine-tuned Language Models
. EMNLP, 2022.
Privacy and Robustness in Federated Learning: Attacks and Defenses
. TNNLS, 2022.
QuoTe: Quality-oriented Testing for Deep Learning Systems
. TOSEM (accepted in 2022).
Machine learning guided alloy design of high-temperature NiTiHf shape memory alloys
. Journal of Materials Science. (2022 Robert W. Cahn Best Paper Award)
2021
Exploring Architectural Ingredients of Adversarially Robust Deep Neural Networks[Code]
. NeurIPS, 2021.
Alpha-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression[Code]
. NeurIPS, 2021.
Anti-Backdoor Learning: Training Clean Models on Poisoned Data[Code]
. NeurIPS, 2021.
Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine Learning
. NeurIPS, 2021.
Unlearnable Examples: Making Personal Data Unexploitable [Code] [Webpage]
. ICLR, 2021. (Spotlight, top 4%) Press: MIT Technology Review, PURSUIT, Gadgets 360.
Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks [Code]
. ICLR, 2021.
Improving Adversarial Robustness via Channel-wise Activation Suppressing [Code]
. ICLR, 2021. (Spotlight, top 4%)
Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better [Code]
. ICCV, 2021.
Noise Doesn’t Lie: Towards Universal Detection of Deep Inpainting
. IJCAI, 2021.
Relationships between Local Intrinsic Dimensionality and Tail Entropy [Video]
. SISAP, Dortmund, Germany, 2021. (Best Paper Award)
RobOT: Robustness-Oriented Testing for Deep Learning Systems [Tookit]
. ICSE, 2021.
Sub-trajectory Similarity Join with Obfuscation
. SSDBM, 2021. (Best Paper Runner-up Award)
SpineOne: A One-Stage Detection Framework for Degenerative Discs and Vertebrae
. BIBM, 2021.
Dual Head Adversarial Training [Code]
. IJCNN, 2021.
Neural Architecture Search via Combinatorial Multi-Armed Bandit
. IJCNN, 2021.
Federated Learning with Extreme Label Skew: A Data Extension Approach
. IJCNN, 2021.
Microwave Link Failures Prediction via LSTM-based Feature Fusion Network
. IJCNN, 2021.
Anomaly Detection for Scenario-based Insider Activities using CGAN Augmented Data
. TrustCom, 2021.
ECG-Adv-GAN: Detecting ECG Adversarial Examples with Conditional Generative Adversarial Networks
. ICMLA, 2021.
Exploring the Vulnerability of Natural Language Processing Models via Universal Adversarial Texts
. ALTA, 2021.
Surgical approach to the facial recess influences the acceptable trajectory of cochlear implantation electrodes
. European Archives of Oto-Rhino-Laryngology, 1-11, 2021.
2020
Normalized Loss Functions for Deep Learning with Noisy Labels [Code]
. ICML, 2020.
Improving Adversarial Robustness Requires Revisiting Misclassified Examples [Code]
. ICLR, 2020.
Skip Connections Matter: on the Transferability of Adversarial Examples Generated with ResNets [Code]
. ICLR, 2020. (Spotlight, top 4%)
Understanding Adversarial Attacks on Deep Learning Based Medical Image Analysis Systems[Code]
. PR, 110, 2021, 107332. (accepted in 2020) Press: Computer Vision News
Clean-Label Backdoor Attacks on Video Recognition Models [Code]
. CVPR, 2020.
Adversarial Camouflage: Hiding Physical-World Attacks with Natural Styles [Code]
. CVPR, 2020.
WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection [Dataset/Code]
. MM, 2020.
Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks [Code]
. ECCV, 2020.
Short-Term and Long-Term Context Aggregation Network for Video Inpainting
. ECCV, 2020. (Spotlight, top 5%)
Transfer of Automated Performance Feedback Models to Different Specimens in Virtual Reality Temporal Bone Surgery
. AIED, 2020.
Towards Fair and Privacy-Preserving Federated Deep Models [Code] [Medium] [Youtube]
. TPDS. (accepted in 2020)
How to Democratise and Protect AI: Fair and Differentially Private Decentralised Deep Learning
. TDSC. (accepted in 2020)
2019
On the Convergence and Robustness of Adversarial Training [Code]
. ICML, Long Beach, USA, 2019. (Long talk, top 3%)
Symmetric Cross Entropy for Robust Learning with Noisy Labels [Code]
. ICCV, Seoul, Korea, 2019.
Black-box Adversarial Attacks on Video Recognition Models [Code]
. MM, Nice, France, 2019.
Generative Image Inpainting with Submanifold Alignment
. IJCAI, Macao, China, 2019.
Exploiting Patterns to Explain Individual Predictions
. KAIS. (accepted in 2019)
2018
Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality [Code]
. ICLR, Vancouver, BC, Canada, 2018, (Oral, top 2%)
Dimensionality-Driven Learning with Noisy Labels [Code]
. ICML, Stockholm, Sweden, 2018. (Long talk, top 4%)
Iterative Learning with Open-set Noisy Labels
. CVPR, Salt Lake City, Utah, USA, 2018. (Spotlight, top 6%)
Providing Automated Real-Time Technical Feedback for Virtual Reality Based Surgical Training: Is the Simpler the Better?
. AIED, London, UK, 2018.
Development and Validation of a Virtual Reality Tutor to Teach Clinically Oriented Surgical Anatomy of the Ear
. CBMS, 2018.
2017
Providing Effective Real-time Feedback in Simulation-based Surgical Training
. MICCAI, Quebec City, Canada, 2017.
Adversarial Generation of Real-time Feedback with Neural Networks for Simulation-based Training
IJCAI, Melbourne, Australia, 2017. (Oral)
No publications match your search. Try another title, author, or keyword.