YuyaoGe's Website
YuyaoGe's Website
About
Coverage
Highlight Papers
Experience
Posts
Projects
Gallery
Light
Dark
Automatic
Home
Tags
Adversarial Attack
Adversarial Attack
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
Prompt-based adversarial attacks have become an effective means to assess the robustness of large language models (LLMs). However, …
Yujia Zheng
,
Tianhao Li
,
Haotian Huang
,
Tianyu Zeng
,
Jingyu Lu
,
Chuangxin Chu
,
Yuekai Huang
,
Ziyou Jiang
,
Qian Xiong
,
Yuyao Ge
,
Mingyang Li
Cite
DOI
PDF
ACL Anthology
arXiv
x
Ask KIMI
Paper Review | Jailbreaking LLMs by Exploiting Decoding Strategies
The authors introduce MaliciousInstruct, a new dataset; a method for evaluating response toxicity; generation exploitation, an attack that manipulates decoding hyperparameters; and generation-aware alignment, a corresponding defence.
Yuyao Ge
Apr 9, 2024
7 min read
Paper Review
Paper
Softmax Regression and Its Optimization
This article is part of my lecture notes from the Deep Learning Systems course taught by Tianqi Chen and J. Zico Kolter at CMU.
Yuyao Ge
Mar 21, 2024
11 min read
Notes
Attack based on data : A novel perspective to attack sensitive points directly
Adversarial attack for time-series classification model is widely explored and many attack methods are proposed. But there is not a …
Yuyao Ge
,
Zhongguo Yang
,
Lizhe Chen
,
Yiming Wang
,
Chengyang Li
Cite
Dataset
PDF
Page
Ask KIMI
Cite
×