A new arXiv paper presents LBA, a textual adversarial attack method for hard-label settings with low query budgets. In this scenario, an attacker can see only the model’s final label, not confidence scores or gradients.
The research matters for robustness testing because many deployed NLP systems expose limited outputs. If effective attacks can be found with few queries, models may be more vulnerable than standard evaluations suggest.
The paper contributes to the broader effort to understand how language systems fail under constrained but realistic adversarial conditions.