2026

NeuroMerge: ML-Guided State Merging for Efficient Symbolic Execution
NeuroMerge: ML-Guided State Merging for Efficient Symbolic Execution

Shenghan Zheng*, Shitong Zhu*, Yu Hao, Xingyu Li, Keyu Man, Zheng Zhang, Qing Deng, Zhiyun Qian, Srikanth V. Krishnamurthy (* equal contribution)

IEEE International Symposium on Software Reliability Engineering (ISSRE) 2026

State merging mitigates path explosion in symbolic execution, but merging the wrong states shifts the cost onto the constraint solver and can slow exploration down instead. NeuroMerge learns when merging pays off, using a machine-learning-guided policy to decide which symbolic states to merge so that exploration stays efficient across large program search spaces.

NeuroMerge: ML-Guided State Merging for Efficient Symbolic Execution

Shenghan Zheng*, Shitong Zhu*, Yu Hao, Xingyu Li, Keyu Man, Zheng Zhang, Qing Deng, Zhiyun Qian, Srikanth V. Krishnamurthy (* equal contribution)

IEEE International Symposium on Software Reliability Engineering (ISSRE) 2026

State merging mitigates path explosion in symbolic execution, but merging the wrong states shifts the cost onto the constraint solver and can slow exploration down instead. NeuroMerge learns when merging pays off, using a machine-learning-guided policy to decide which symbolic states to merge so that exploration stays efficient across large program search spaces.

ClawsBench: High Fidelity Simulated Workspace Environments for Evaluating and Improving Productivity Agents
ClawsBench: High Fidelity Simulated Workspace Environments for Evaluating and Improving Productivity Agents

Xiangyi Li, Kyoung Whan Choe, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Shenghan Zheng, Xiaokun Chen, Chujun Tao, Weixiang Yan, Jiankai Sun, Yiyuan Li, Jiajun Bao, Yuanli Wang, Hanchung Lee

Conference on Language Modeling (COLM) 2026

Evaluating LLM agents on live productivity services risks irreversible errors and lacks reproducibility. We introduce smolclaws, a framework of five high-fidelity simulated services (Gmail, Google Calendar, Google Docs, Google Drive, and Slack) with full state management and deterministic replay, paired with 40+ tasks spanning multi-step workflows, cross-service coordination, and safety-critical scenarios. Across 1,500+ trials frontier agents solve only 30% of tasks, while trajectory-driven skill refinement yields +35% average reward.

ClawsBench: High Fidelity Simulated Workspace Environments for Evaluating and Improving Productivity Agents

Xiangyi Li, Kyoung Whan Choe, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Shenghan Zheng, Xiaokun Chen, Chujun Tao, Weixiang Yan, Jiankai Sun, Yiyuan Li, Jiajun Bao, Yuanli Wang, Hanchung Lee

Conference on Language Modeling (COLM) 2026

Evaluating LLM agents on live productivity services risks irreversible errors and lacks reproducibility. We introduce smolclaws, a framework of five high-fidelity simulated services (Gmail, Google Calendar, Google Docs, Google Drive, and Slack) with full state management and deterministic replay, paired with 40+ tasks spanning multi-step workflows, cross-service coordination, and safety-critical scenarios. Across 1,500+ trials frontier agents solve only 30% of tasks, while trajectory-driven skill refinement yields +35% average reward.

NCFuzz: Configuration-guided Network Service Fuzzing
NCFuzz: Configuration-guided Network Service Fuzzing

Xuesong Bai, Hengkai Ye, Shenghan Zheng, Fenglu Zhang, Hong Hu, Zhou Li

ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026

Network services expose much of their behavior only under particular configurations, so a fuzzer that leaves the configuration space untouched never reaches large parts of the code. NCFuzz treats configuration as a first-class input dimension, using it to guide network service fuzzing toward states and code paths that conventional input-only fuzzing does not reach.

NCFuzz: Configuration-guided Network Service Fuzzing

Xuesong Bai, Hengkai Ye, Shenghan Zheng, Fenglu Zhang, Hong Hu, Zhou Li

ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026

Network services expose much of their behavior only under particular configurations, so a fuzzer that leaves the configuration space untouched never reaches large parts of the code. NCFuzz treats configuration as a first-class input dimension, using it to guide network service fuzzing toward states and code paths that conventional input-only fuzzing does not reach.

Formal Security Analysis of Agent Protocol Composition
Formal Security Analysis of Agent Protocol Composition

Shenghan Zheng, Qifan Zhang, Zheng Zhang, Haonan Li, Christophe Hauser

arXiv preprint 2026

AI agent protocols define how agents use tools, delegate work, and coordinate across software systems, but their security requirements remain incomplete and inconsistently enforced. We present AgentThread, a source-linked framework that analyzes agent protocols from specification text to running SDKs, formalizing protocol-derived checks as TLA+ invariants and replaying executable counterexamples against real SDKs. Across five emerging protocols it identifies 35 specification-level findings and 30 further failures that emerge only under protocol composition.

Formal Security Analysis of Agent Protocol Composition

Shenghan Zheng, Qifan Zhang, Zheng Zhang, Haonan Li, Christophe Hauser

arXiv preprint 2026

AI agent protocols define how agents use tools, delegate work, and coordinate across software systems, but their security requirements remain incomplete and inconsistently enforced. We present AgentThread, a source-linked framework that analyzes agent protocols from specification text to running SDKs, formalizing protocol-derived checks as TLA+ invariants and replaying executable counterexamples against real SDKs. Across five emerging protocols it identifies 35 specification-level findings and 30 further failures that emerge only under protocol composition.

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Yifeng He, Shenghan Zheng, et al.

arXiv preprint 2026

Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Curated Skills raise the average pass rate from 33.9% to 50.5% across 18 model-harness configurations, and focused Skills with at most three modules outperform larger, exhaustive bundles.

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Yifeng He, Shenghan Zheng, et al.

arXiv preprint 2026

Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Curated Skills raise the average pass rate from 33.9% to 50.5% across 18 model-harness configurations, and focused Skills with at most three modules outperform larger, exhaustive bundles.

2025

SCAD: Towards a Universal and Automated Network Side-Channel Vulnerability Detection
SCAD: Towards a Universal and Automated Network Side-Channel Vulnerability Detection

Keyu Man, Zhongjie Wang, Yu Hao, Shenghan Zheng, Yue Cao, Xin'an Zhou, Zhiyun Qian

IEEE Symposium on Security and Privacy (IEEE S&P) 2025

Network side-channel attacks, such as SADDNS enabling off-path cache poisoning, are notoriously difficult to detect because current automated techniques require extensive, error-prone modeling that oversimplifies network protocols. In response, we introduce SCAD—the first solution leveraging dynamic symbolic execution to efficiently identify non-interference violations across multiple execution traces—uncovering previously unknown vulnerabilities with significantly reduced manual effort.

SCAD: Towards a Universal and Automated Network Side-Channel Vulnerability Detection

Keyu Man, Zhongjie Wang, Yu Hao, Shenghan Zheng, Yue Cao, Xin'an Zhou, Zhiyun Qian

IEEE Symposium on Security and Privacy (IEEE S&P) 2025

Network side-channel attacks, such as SADDNS enabling off-path cache poisoning, are notoriously difficult to detect because current automated techniques require extensive, error-prone modeling that oversimplifies network protocols. In response, we introduce SCAD—the first solution leveraging dynamic symbolic execution to efficiently identify non-interference violations across multiple execution traces—uncovering previously unknown vulnerabilities with significantly reduced manual effort.