Research Project Pages

Indexed project pages for published papers — summaries, figures, BibTeX, and outbound links for discovery.

CogniDir

·EMNLP·2026

LLM-generated cognitive malicious comments exploit fact distortion, logical confusion, and emotional manipulation to break comment-aware fake news detectors. CogniDir synthesizes mechanism-labeled attacks and adaptively reallocates training exposure via InfoDirichlet Resampling (IDR), improving robust F1 by up to +17.9% on three cross-lingual benchmarks.

CoT is Not the Chain of Truth

·ICML·2026

Reasoning LLMs can harbor unsafe planning inside Chain-of-Thought (CoT) traces even when the final answer refuses. We localize these failures to mid-depth layers and safety-critical attention heads via Jacobian spectral metrics, validate the pattern beyond fake-news generation on HarmBench, and show parameter-efficient mitigation.