Change search
Link to record
Permanent link

Direct link
Publications (5 of 5) Show all publications
Fu, J., Wu, Y., Chen, Y., Peng, K., Zhang, X., Cevher, V., . . . Holst, A. (2026). Diffusion-based Cumulative Adversarial Purification for Vision Language Models. Transactions on Machine Learning Research, 2026-June
Open this publication in new window or tab >>Diffusion-based Cumulative Adversarial Purification for Vision Language Models
Show others...
2026 (English)In: Transactions on Machine Learning Research, E-ISSN 2835-8856, Vol. 2026-JuneArticle in journal (Refereed) Published
Abstract [en]

Vision Language Models (VLMs) have shown remarkable capabilities in multimodal under-standing, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications. Despite often being imperceptible to humans, these perturbations can drastically alter model outputs, leading to erroneous interpretations and decisions. This paper introduces DiffCAP, a novel diffusion-based purification strategy that can effectively neutralize adversarial corruptions in VLMs. We theoretically establish a provable recovery region in the forward diffusion process and meanwhile quantify the convergence rate of semantic variation with respect to VLMs. These findings manifest that adversarial effects monotonically fade as diffusion unfolds. Guided by this principle, DiffCAP leverages noise injection with a similarity threshold of VLM embeddings as an adaptive criterion, before reverse diffusion restores a clean and reliable representation for VLM inference. Through extensive experiments across six datasets with three VLMs under varying attack strengths in three task scenarios, we show that DiffCAP outperforms existing defense techniques by a substantial margin. Notably, DiffCAP significantly reduces both hyperparameter tuning complexity and the required diffusion time, thereby accelerating the denoising process. Equipped with theorems and empirical support, DiffCAP provides a robust and practical solution for securely deploying VLMs in adversarial environments. The source code is available at https://github.com/JasonFu1998/DiffCAP

Place, publisher, year, edition, pages
Transactions on Machine Learning Research, 2026
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:ri:diva-81839 (URN)2-s2.0-105041547155 (Scopus ID)
Note

QC 20260625

Available from: 2026-06-25 Created: 2026-06-25 Last updated: 2026-06-25Bibliographically approved
Huang, W., Euler, N., Lloyd, K. A., Diaz Boada, J. S., Fu, J., Gunasekera, S., . . . Grönwall, C. (2026). Variable fab domain N-glycosylation patterns in the B cell receptor repertoires of healthy individuals and patients with rheumatoid arthritis. Journal of Immunology, 215(6)
Open this publication in new window or tab >>Variable fab domain N-glycosylation patterns in the B cell receptor repertoires of healthy individuals and patients with rheumatoid arthritis
Show others...
2026 (English)In: Journal of Immunology, ISSN 0022-1767, E-ISSN 1550-6606, Vol. 215, no 6Article in journal (Refereed) Published
Abstract [en]

N-linked glycosylation (N-glyc) sites (N-X-S/T, X≠P) can be introduced by somatic hypermutation in immunoglobulin Fab regions. In patients with rheumatoid arthritis (RA), anti-citrullinated protein antibodies have a striking overrepresentation of Fab N-glycosylation. To further explore this, we sequenced B cell receptors (BCRs) from peripheral blood of 13 RA patients and 6 healthy control subjects, analyzing in total >250,000 heavy chain (VH) and >100,000 light chain sequences from both total B cells and citrullinated fibrinogen-reactive (Cit-Fib+) cells. Distribution of variable VH genes in VDJ DNA, and transcripts of unmutated IgM and class-switched BCR, revealed transcript gene-usage bias and higher VH4 in natural rearrangements by out-of-frame VDJ DNA in RA patients compared with control subjects. IgG Fab N-glyc sites were slightly more prominent in RA than control subjects (14.9% versus 12.1%; P = 0.048) with certain VH genes (e.g. VH1-18, VH1-69, VH3-9) displaying enriched N-glyc cumulative frequencies by somatic hypermutation. VH gene N-glyc hotspots were identified, explained by a lower threshold for codon conversion (i.e. K/S/T-X-S/T) and especially frequent in VH4s. Yet, RA patients had significantly more N-glyc in complementarity-determining region (CDR) 1 and 3 compared with control subjects. Expanded clonotypes with somatic hypermutation-induced N-glyc sites were delineated by network analysis in both the RA patients and control group, but patients with RA displayed more highly mutated class-switched members. Furthermore, Cit-Fib+ BCR-expanded clonotypes could be traced in the total B cell repertoire and exhibited increased frequency of N-glyc in mutated IgG/IgA subsets. Our findings highlight how Fab N-glyc sites can be linked to biased clonotype evolution and B cell selection in chronic responses

Place, publisher, year, edition, pages
Oxford University Press (OUP), 2026
Keywords
BCR repertoire, glycosylation, immunoglobulin, rheumatoid arthritis, V-gene sequencing
National Category
Immunology in the Medical Area
Identifiers
urn:nbn:se:ri:diva-81940 (URN)10.1093/jimmun/vkag113 (DOI)42269014 (PubMedID)2-s2.0-105042210463 (Scopus ID)
Note

QC 20260713

Available from: 2026-07-13 Created: 2026-07-13 Last updated: 2026-07-13Bibliographically approved
Fu, J., Zhang, X., Pashami, S., Rahimian, F. & Holst, A. (2025). DiffPAD: Denoising Diffusion-Based Adversarial Patch Decontamination. In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV): . Paper presented at 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 6602-6611). Institute of Electrical and Electronics Engineers Inc.
Open this publication in new window or tab >>DiffPAD: Denoising Diffusion-Based Adversarial Patch Decontamination
Show others...
2025 (English)In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Institute of Electrical and Electronics Engineers Inc. , 2025, p. 6602-6611Conference paper, Published paper (Refereed)
Abstract [en]

In the ever-evolving adversarial machine learning landscape, developing effective defenses against patch attacks has become a critical challenge, necessitating reliable solutions to safeguard real-world AI systems. Although diffusion models have shown remarkable capacity in image synthesis and have been recently utilized to counter lp-norm bounded attacks, their potential in mitigating localized patch attacks remains largely underexplored. In this work, we propose DiffPAD, a novel framework that harnesses the power of diffusion models for adversarial patch decontamination. DiffPAD first performs super-resolution restoration on downsampled input images, then adopts binarization, dynamic thresholding scheme and sliding window for effective localization of adversarial patches. Such a design is inspired by the theoretically derived correlation between patch size and diffusion restoration error that is generalized across diverse patch attack scenarios. Finally, DiffPAD applies inpainting techniques to the original input images with the estimated patch region being masked. By integrating closed-form solutions for super-resolution restoration and image inpainting into the conditional reverse sampling process of a pre-trained diffusion model, DiffPAD obviates the need for text guidance or fine-tuning. Through comprehensive experiments, we demonstrate that DiffPAD not only achieves state-of-the-art adversarial robustness against patch attacks but also excels in recovering naturalistic images without patch remnants. The source code is available at https://github.com/JasonFu1998/DiffPAD. 

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers Inc., 2025
Keywords
Adversarial machine learning; Image coding; Photointerpretation; Adversarial defense; AI systems; Critical challenges; De-noising; Diffusion model; Input image; Machine-learning; Patch attack; Real-world; Super-resolution restoration; Decontamination
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-78562 (URN)10.1109/WACV61041.2025.00643 (DOI)2-s2.0-105003628690 (Scopus ID)
Conference
2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Available from: 2025-09-16 Created: 2025-09-16 Last updated: 2026-01-22Bibliographically approved
Peng, K., Fu, J., Yang, K., Wen, D., Chen, Y., Liu, R., . . . Roitberg, A. (2025). Referring Atomic Video Action Recognition. Paper presented at 18th European Conference on Computer Vision, ECCV 2024. Milan, Italy. 29 September 2024 through 4 October 202. Lecture Notes in Computer Science, 15077 LNCS, 166-185
Open this publication in new window or tab >>Referring Atomic Video Action Recognition
Show others...
2025 (English)In: Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349, Vol. 15077 LNCS, p. 166-185Article in journal (Refereed) Published
Abstract [en]

We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from traditional action recognition and localization, where predictions are delivered for all present individuals. In contrast, we focus on recognizing the correct atomic action of a specific individual, guided by text. To explore this task, we present the RefAVA dataset, containing 36, 630 instances with manually annotated textual descriptions of the individuals. To establish a strong initial benchmark, we implement and validate baselines from various domains, e.g., atomic action localization, video question answering, and text-video retrieval. Since these existing methods underperform on RAVAR, we introduce RefAtomNet – a novel cross-stream attention-driven method specialized for the unique challenges of RAVAR: the need to interpret a textual referring expression for the targeted individual, utilize this reference to guide the spatial localization and harvest the prediction of the atomic actions for the referring person. The key ingredients are: (1) a multi-stream architecture that connects video, text, and a new location-semantic stream, and (2) cross-stream agent attention fusion and agent token fusion which amplify the most relevant information across these streams and consistently surpasses standard attention-based fusion on RAVAR. Extensive experiments demonstrate the effectiveness of RefAtomNet and its building blocks for recognizing the action of the described individual. The dataset and code will be made publicly available at RAVAR.

Place, publisher, year, edition, pages
Springer Science and Business Media Deutschland GmbH, 2025
Keywords
Action recognition; Atomic actions; Localisation; Multi-stream architecture; Question Answering; Referring expressions; Spatial localization; Textual description; Video data; Video retrieval; Semantics
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-77988 (URN)10.1007/978-3-031-72655-2_10 (DOI)2-s2.0-85213009172 (Scopus ID)
Conference
18th European Conference on Computer Vision, ECCV 2024. Milan, Italy. 29 September 2024 through 4 October 202
Note

The project served to prepare the SFB 1574 Circular Factory for the Perpetual Product (project ID: 471687386), approved by the German Research Foundation (DFG, German Research Foundation) with a start date of April 1, 2024. This work was also partially supported in part by the SmartAge project sponsored by the Carl Zeiss Stiftung (P2019-01-003; 2021\u20132026). This work was performed on the HoreKa supercomputer funded by the Ministry of Science, Research and the Arts Baden-W\"urttemberg and by the Federal Ministry of Education and Research. The authors also acknowledge support by the state of Baden-W\"urttemberg through bwHPC and the German Research Foundation (DFG) through grant INST 35/1597-1 FUGG. This project is also supported by the National Key RD Program under Grant 2022YFB4701400. Lastly, the authors thank for the support of Dr. Sepideh Pashami, the Swedish Innovation Agency VINNOVA, the Digital Futures.

Available from: 2025-03-03 Created: 2025-03-03 Last updated: 2025-09-23Bibliographically approved
Fu, J., Tan, J., Yin, W., Pashami, S. & Björkman, M. (2023). Component attention network for multimodal dance improvisation recognition. In: : . Paper presented at 25th International Conference on Multimodal Interaction, ICMI 2023. Paris, France. 9 October 2023 through 13 October 2023 (pp. 114-118). Association for Computing Machinery
Open this publication in new window or tab >>Component attention network for multimodal dance improvisation recognition
Show others...
2023 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Dance improvisation is an active research topic in the arts. Motion analysis of improvised dance can be challenging due to its unique dynamics. Data-driven dance motion analysis, including recognition and generation, is often limited to skeletal data. However, data of other modalities, such as audio, can be recorded and benefit downstream tasks. This paper explores the application and performance of multimodal fusion methods for human motion recognition in the context of dance improvisation. We propose an attention-based model, component attention network (CANet), for multimodal fusion on three levels: 1) feature fusion with CANet, 2) model fusion with CANet and graph convolutional network (GCN), and 3) late fusion with a voting strategy. We conduct thorough experiments to analyze the impact of each modality in different fusion methods and distinguish critical temporal or component features. We show that our proposed model outperforms the two baseline methods, demonstrating its potential for analyzing improvisation in dance

Place, publisher, year, edition, pages
Association for Computing Machinery, 2023
Keywords
Arts computing; Attention network; Dance recognition; Data driven; Down-stream; Fusion methods; Improvization; Multi-modal; Multi-modal fusion; Performance; Research topics; Motion estimation
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:ri:diva-67967 (URN)10.1145/3577190.3614114 (DOI)2-s2.0-85175844284 (Scopus ID)
Conference
25th International Conference on Multimodal Interaction, ICMI 2023. Paris, France. 9 October 2023 through 13 October 2023
Available from: 2023-11-24 Created: 2023-11-24 Last updated: 2025-09-23Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0009-0004-3798-8603

Search in DiVA

Show all publications