Discovery window: 2026-08-01 – 2026-09-16
Candidates checked: 17 · New works: 2 · Known unchanged: 13 · Version/publication updates: 0
Selected new works: 2 · P0: 2 · P1: 0 · P2: 0
Focus: AI4SE first; general AI only when it can materially affect software-engineering research or practice.

0. Deduplication summary

  • New works: 2
  • Suppressed as already known and unchanged: 13
  • Known works with meaningful updates: 0
  • Possible duplicates requiring identity check: 0

First-pass deduplication used papers/registry/index.json schema v2. Exact arXiv IDs were checked before title/author evidence. The two selected works did not occur in the identifier or title index.


1. Today's signal

What is worth noticing today?

  • Coding-agent benchmarks are moving from isolated issue fixing toward realistic autonomy boundaries. SWE-Touch adds concurrent user edits to the workspace; Active-SWE removes the issue report and asks the agent to discover defects itself.
  • Repository state awareness is now measurable rather than anecdotal. SWE-Touch shows that a small, plausible user edit can invalidate an agent's latent model of the repository and reduce resolution performance even when the underlying issue is unchanged.
  • Bug discovery remains much harder than bug diagnosis. Active-SWE's 1,663 tasks expose the gap between being told what to fix and autonomously deciding what is wrong, where it is, and whether a proposed new bug is valid.

Must-read shortlist

PriorityPaperAreaStatus / VenueWhy it mattersAction
P0SWE-Touchcoding-agents human-AI collaborationPreprint — venue not statedDirectly tests coding agents in shared workspaces where users modify code during execution.Read / Queue
P0Active-SWEcoding-agents APR bug discoveryPreprint — venue not statedRemoves issue-report guidance and evaluates proactive single/multi-bug discovery and repair at large scale.Read / Queue

2. New priority papers

[P0] SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

  • Work ID: arxiv-2608.02499
  • Registry: record
  • Links: Paper · PDF source · Code
  • Primary area: coding agents / human-AI collaboration
  • Tags: SWE-bench shared-workspace state-awareness
  • Status: Preprint — venue not stated
  • Venue: arXiv
  • Version / date: arXiv v1, 2026-08-03
  • Authors: Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu
  • Affiliations / team: Institute of Automation, Chinese Academy of Sciences / University of Chinese Academy of Sciences
  • Notable author/team signal: CAS NLP/agent team; open implementation released.
  • First seen: 2026-09-16

Fast grasp

One-sentence takeaway.
SWE-Touch injects validated task-conflicting user edits into live coding-agent trajectories and finds an average 7.7 percentage-point resolve-rate drop on SWE-bench Verified, showing that strong autonomous coding does not imply robust collaboration in a mutable shared workspace.

Research problem.
How well do coding agents detect, reconcile and validate code changes made by a user while the agent is already working on the same repository?

Problem definition / setting.
An agent works on a repository-level repair task while a user can alter task-relevant code in the same workspace. The benchmark injects a plausible Counter-Edit at a trajectory-relevant point, accompanied by contextual user messaging, and measures whether the agent adapts rather than proceeding from stale repository assumptions.

Core idea.
Treat the repository itself as a second human-agent communication channel. Instead of testing only textual interaction, perturb task-critical code during execution and measure state awareness and recovery.

Method / system.

  • Mines task-critical regions from multiple repair trajectories.
  • Uses a separate User Patch Generator to construct plausible Counter-Edits that conflict with task completion.
  • Injects edits when agents reach relevant code and supplies contextual user messages.
  • Evaluates resulting trajectories for reinspection, reconciliation and targeted validation behavior.

Evaluation.

  • Benchmarks / datasets: SWE-bench Verified; additional longer-horizon experiments on SWE-Bench Pro and DeepSWE.
  • Baselines: nine coding models/agents.
  • Metrics: task resolve rate and trajectory-level behavioral analysis.
  • Scale / setup: repository-level interactive perturbation during agent execution.

Main findings.

  • Counter-Edits reduce average resolve rate by 7.7 percentage points on SWE-bench Verified.
  • Degradation persists on longer-horizon SWE-Bench Pro and DeepSWE tasks.
  • Failures commonly involve retaining conflicting code or overwriting it without sufficient repository reinspection and targeted tests.

Novelty vs. prior work.
The novelty is the interaction channel: prior repository benchmarks largely assume an agent-owned workspace, and human-in-the-loop benchmarks usually restrict users to messages. SWE-Touch makes concurrent code edits a controlled benchmark variable.

Why it matters for AI4SE.
This is directly relevant to real IDE/agent deployment. Repository-level agents need explicit change detection, workspace synchronization and conflict-aware validation; otherwise benchmark autonomy can overstate practical pair-programming reliability.

Limitations / concerns.

  • Counter-Edits are generated perturbations and may not fully represent the intent distribution of organic user edits.
  • Results depend on when and where edits are injected; benchmark policy can influence measured robustness.
  • The benchmark measures adversarial/conflicting edits more strongly than cooperative user contributions.

Reading recommendation.
Read now — prioritize Counter-Edit construction, injection policy and trajectory failure analysis; these are reusable for human-agent and mutable-environment benchmark design.


[P0] Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

  • Work ID: arxiv-2608.04682
  • Registry: record
  • Links: Paper · PDF source · Code
  • Primary area: coding agents / APR / bug detection
  • Tags: proactive-repair fault-localization benchmark
  • Status: Preprint — venue not stated
  • Venue: arXiv
  • Version / date: arXiv v1, 2026-08-05
  • Authors: Haobin Li, Ping Deng, Weizhong Qian, Liang Jiang, Zhenyu Huang, Mouxing Yang, Xi Peng
  • Affiliations / team: not conservatively asserted from the sources checked
  • Notable author/team signal: public benchmark implementation and evaluation artifacts released by XLearning-SCU.
  • First seen: 2026-09-16

Fast grasp

One-sentence takeaway.
Active-SWE builds 1,663 tasks across six bug categories and eight languages to test whether coding agents can discover and fix bugs without an issue report; current agents remain weak at localization, multi-bug repair and validating newly discovered defects.

Research problem.
Can coding agents proactively audit a repository, identify defects and repair them when no human-authored issue report tells them what to look for?

Problem definition / setting.
The agent receives repository context and a generic senior-engineer audit instruction rather than a specific issue description. Tasks include recorded bugs, multi-bug combinations and potential-bug discovery, requiring localization and repair decisions before patch generation.

Core idea.
Remove the issue-report oracle that standard SWE-bench-style evaluation provides, then separately measure recovery of known defects and evidence-backed discovery of additional valid defects.

Method / system.

  • Mines high-quality pull requests from 87 popular open-source repositories.
  • Uses a six-category functional bug taxonomy and covers eight programming languages.
  • Builds Docker execution environments with automated setup and distinguishing tests.
  • Creates simple single-bug and harder temporally composed multi-bug settings.
  • Uses a dual-track evaluation for recorded-bug repair and newly revealed bugs.

Evaluation.

  • Benchmarks / datasets: Active-SWE, 1,663 tasks; six bug categories; eight languages.
  • Baselines: multiple state-of-the-art coding agents.
  • Metrics: localization recall/precision, Resolved, generated-test Count, test validity and evidence-backed Revealed bugs.
  • Scale / setup: tasks mined from 87 open-source repositories with containerized execution.

Main findings.

  • State-of-the-art agents struggle to locate and resolve recorded bugs without issue-report guidance.
  • Multi-bug settings substantially increase difficulty.
  • Potential-bug discovery is especially demanding because agents must provide executable evidence rather than merely flag suspicious code.

Novelty vs. prior work.
The important shift is from reactive issue resolution to proactive repository auditing. The benchmark also distinguishes fixing known historical defects from discovering additional defects, which broadens coding-agent evaluation toward autonomous maintenance.

Why it matters for AI4SE.
Active-SWE joins fault localization, APR and coding agents in one setting. It tests a capability needed for autonomous maintenance systems but mostly absent from issue-driven benchmarks: deciding what deserves repair before deciding how to repair it.

Limitations / concerns.

  • Mining from historical PRs still gives the benchmark a retrospective ground truth and may not fully model open-ended production auditing.
  • Generated/setup-agent infrastructure and judge validation can introduce construction bias.
  • Proactive discovery metrics depend on what counts as sufficient executable evidence for a newly claimed bug.

Reading recommendation.
Read now — focus on task construction, the dual-track evaluator and the gap between localization and successful resolution; compare directly with SWE-bench and fault-localization benchmarks.


3. Other new selected papers

No P2 papers selected today.


4. Publication and version updates for known works

No meaningful publication/version updates were confirmed for already indexed works today.


5. Possible duplicates requiring review

None.


6. Reading-queue actions

Add

  • arxiv-2608.02499 — SWE-Touch — high-value benchmark for shared-workspace human-agent collaboration — priority P0
  • arxiv-2608.04682 — Active-SWE — high-value proactive bug discovery/repair benchmark — priority P0

Promote to deep read

  • arxiv-2608.02499 — benchmark perturbation design is directly useful for coding-agent reliability research.
  • arxiv-2608.04682 — proactive repair formulation connects fault localization, APR and autonomous maintenance.

Remove / deprioritize

  • None.

7. Coverage and search notes


Sources checked

Software Engineering journals / venues

  • [x] TSE
  • [x] TOSEM
  • [x] EMSE
  • [x] JSS
  • [x] ICSE
  • [x] FSE / ESEC-FSE
  • [x] ASE
  • [x] ISSTA
  • [x] MSR
  • [x] ACM SIGSOFT publications / relevant SIGSOFT venues

General AI / ML / NLP venues

  • [x] ICML
  • [x] NeurIPS
  • [x] ICLR
  • [x] ACL
  • [x] AAAI

arXiv

  • [x] cs.SE
  • [x] cs.AI intersecting software engineering
  • [x] cs.CL intersecting software engineering
  • [x] cs.LG intersecting software engineering

High-value signals

  • [x] Notable researchers / labs
  • [x] Major technology companies
  • [x] New benchmark or dataset releases
  • [x] New code/model releases attached to recent papers

Coverage gaps / failures: no blocking source failures. Same-day venue indexing can lag arXiv and author/project pages, so publication status remains conservative.


8. Notes

Together with the recently tracked SWE-bench Science and SWE Refactor Bench, today's papers reinforce a clear evaluation trend: the next useful frontier is not simply harder issue fixing, but testing agents under missing specifications, mutable workspaces, multiple simultaneous defects and stronger evidence requirements.