Research Project Pages
Indexed project pages for published papers — summaries, figures, BibTeX, and outbound links for discovery.
·EMNLP·2026
LLM-generated cognitive malicious comments exploit fact distortion, logical confusion, and emotional manipulation to break comment-aware fake news detectors. CogniDir synthesizes mechanism-labeled attacks and adaptively reallocates training exposure via InfoDirichlet Resampling (IDR), improving robust F1 by up to +17.9% on three cross-lingual benchmarks.
Project page →·arXiv·PDF
·ICML·2026
Reasoning LLMs can harbor unsafe planning inside Chain-of-Thought (CoT) traces even when the final answer refuses. We localize these failures to mid-depth layers and safety-critical attention heads via Jacobian spectral metrics, validate the pattern beyond fake-news generation on HarmBench, and show parameter-efficient mitigation.
Project page →·arXiv·PDF