Discovery window: 2026-09-15 – 2026-09-18
Candidates checked: 19 · New works: 3 · Known unchanged: 16 · Version/publication updates: 0
Selected new works: 3 · P0: 2 · P1: 1 · P2: 0
Focus: AI4SE first; general AI only when it can materially affect software-engineering research or practice.

0. Deduplication summary

  • New works: 3
  • Suppressed as already known and unchanged: 16
  • Known works with meaningful updates: 0
  • Possible duplicates requiring identity check: 0

First-pass deduplication used papers/registry/index.json (schema v2, 16 works). Exact stable identifiers were checked before title/author evidence. The three selected arXiv identifiers were absent from the identifier index and their normalized titles did not collide with indexed works.


1. Today's signal

What is worth noticing today?

  • Coding-agent evaluation is moving away from a single leaderboard number. The SWE-bench audit shows that adjacent frontier entries can be statistically indistinguishable even when the leaderboard presents a strict ordering; model–scaffold provenance and paired outcome structure matter.
  • Specifications are becoming executable and interactive. ProgramDistill replaces a textual issue with a working reference application: the agent must discover behavior through interaction and recreate it, while replayable traces provide scalable verification.
  • Repository context is becoming adaptive and multimodal. RepoAtlas treats repository understanding as a stateful select–project–refresh loop rather than a one-shot retrieval step, improving both task success and context efficiency.

Must-read shortlist

PriorityPaperAreaStatus / VenueWhy it mattersAction
P0ProgramDistillcoding-agents benchmarkPreprint — venue not statedIntroduces reference-guided SWE tasks with 4,063 automatically constructed, replay-verifiable tasks.Read / Queue
P0Coding Agents Have Convergedcoding-agents evaluationAccepted — ADMA 2026Shows current SWE-bench Verified top-rank gaps often lack statistical resolution and quantifies scaffold dependence.Read / Queue
P1RepoAtlasrepository-navigation multimodalPreprint — venue not statedTraining-free evolving code-graph views improve SWE-bench Verified while reducing tokens and calls.Queue

2. New priority papers

[P0] ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

  • Work ID: arxiv-2609.18805
  • Registry: record
  • Links: Paper · PDF source · Project
  • Primary area: coding agents / repository-level web development
  • Tags: benchmark reference-guided-SWE executable-specification
  • Status: Preprint — venue not stated
  • Venue: arXiv
  • Version / date: arXiv v1, 2026-09-16
  • Authors: Jeonghye Kim, Minseon Kim, Young Jin Kim, Matheus Pereira, Marc-Alexandre Côté, Alessandro Sordoni, Xingdi Yuan, Zhengyan Shi
  • Affiliations / team: Microsoft Research Montréal / KAIST collaboration
  • Notable author/team signal: Microsoft Research team; project is presented through the Debug-Gym research site.
  • First seen: 2026-09-18

Fast grasp

One-sentence takeaway.
ProgramDistill creates 4,063 coding-agent tasks from 1,975 replay-verified behaviors across 26 web applications, evaluating whether agents can infer missing functionality by interacting with a working reference rather than reading a complete textual specification.

Research problem.
How should coding agents be evaluated when intended behavior is demonstrated by existing software rather than fully specified in an issue, test, or natural-language requirement?

Problem definition / setting.
The agent receives an incomplete editable web application plus access to a fully functional reference application. It must explore the reference through browser interaction, infer the target behavior, modify the incomplete implementation, and satisfy executable replay checks without access to the reference source or gold patch.

Core idea.
Turn working applications into scalable executable specifications. The mine-craft-patch pipeline discovers reproducible behaviors, removes their implementations while preserving prerequisites, and uses the same replay trace to validate the original, masked, and repaired application.

Method / system.

  • Mining agents explore reference applications and record browser-action traces with expected success signals and prerequisite links.
  • Crafting agents remove implementations for selected behaviors; build/replay checks ensure prerequisites remain valid and the target behavior becomes broken.
  • Gold patches verify recoverability; coding agents then repair only through interaction with the reference application.
  • Feature dependencies provide a natural restoration-depth axis for controlled difficulty and potential curriculum learning.

Evaluation.

  • Benchmarks / datasets: ProgramDistill, 26 applications, 1,975 replay-verified behaviors, 4,063 tasks.
  • Baselines: nine frontier coding agents.
  • Metrics: replay-based task success; cumulative workflow success; success versus restoration depth.
  • Scale / setup: full-application and partial-application reconstruction with executable browser traces.

Main findings.

  • In full-application cumulative workflows, GPT-6 Astra reaches 49.2% success and Claude Opus 5 reaches 28.8%.
  • In partial reconstruction, increasing restoration depth from 1 to 8 drops success from 100% to 64.0% for one reported agent and from 96% to 32% for another.
  • The pipeline constructs thousands of verifiable tasks without human task authoring, suggesting a scalable route to both evaluation and training data.

Novelty vs. prior work.
The key novelty is the specification channel: desired behavior is discovered by interacting with working software, not supplied as an issue or frozen test description. Replay traces simultaneously provide behavioral evidence, dependency structure, and an automatic verifier.

Why it matters for AI4SE.
This setting is close to maintenance, reimplementation, compatibility engineering, UI regression repair, and reverse engineering. It also offers a benchmark-construction pattern that is less dependent on historical GitHub issues and naturally supports curriculum/RL data generation.

Limitations / concerns.

  • Current tasks are web applications, so generalization to backend, systems, mobile, or scientific software is untested.
  • Replayable browser behavior captures functional interaction but may miss visual fidelity, nonfunctional requirements, security, accessibility, or hidden state.
  • The benchmark-generation pipeline itself uses multiple LLM agents, so task diversity and failure modes may reflect generator-model biases.

Reading recommendation.
Read now — prioritize the mine-craft-patch validity checks, dependency/restoration-depth construction, and cumulative-workflow evaluation; these are the most reusable ideas for new coding-agent benchmarks.


[P0] Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

  • Work ID: arxiv-2609.17394
  • Registry: record
  • Links: Paper · PDF source · Code
  • Primary area: coding-agent evaluation / empirical AI4SE
  • Tags: SWE-bench statistical-power model-scaffold
  • Status: Accepted — ADMA 2026, Special Session on Responsible Data Intelligence
  • Venue: ADMA 2026
  • Version / date: arXiv v1 camera-ready, 2026-09-15
  • Authors: Fengshuo Liu, Ying Liu, Ruize Sun, Lie Luo, Siyuan Guo
  • Affiliations / team: not asserted in registry
  • Notable author/team signal: —
  • First seen: 2026-09-18

Fast grasp

One-sentence takeaway.
An audit of 254 SWE-bench submissions finds that the top of SWE-bench Verified has insufficient statistical resolution for the strict rankings commonly inferred from small score gaps, while scaffold choice can move the same model far more than the spread among top leaderboard entries.

Research problem.
Do small aggregate differences on SWE-bench leaderboards provide enough statistical evidence to rank frontier coding-agent systems, and how much of observed performance is attributable to the model versus its scaffold?

Problem definition / setting.
Rather than rerunning models, the study analyzes per-instance outcomes from 254 published submissions across four SWE-bench splits. It measures overlap/nesting of solved sets, paired statistical distinguishability, model–scaffold variation, grouping sensitivity, and the instance budget required to resolve differences.

Core idea.
Treat coding-agent leaderboard comparison as a paired statistical-resolution problem over task outcomes, not as an ordering induced automatically by aggregate percentages.

Method / system.

  • Audits shared successes/failures and computes nesting of solution sets.
  • Uses exact paired McNemar tests for adjacent leaderboard entries.
  • Separates model and scaffold observations and tests model–scaffold interactions.
  • Provides a five-step audit protocol covering shared outcomes, paired differences, grouping sensitivity, and required instance budget.

Evaluation.

  • Benchmarks / datasets: 254 SWE-bench submissions across four splits, including Verified and Test.
  • Baselines: published leaderboard submissions rather than new agent runs.
  • Metrics: resolved counts, solution-set nesting, exact paired McNemar tests, within-model scaffold range, grouping/tier sensitivity.
  • Scale / setup: SWE-bench Verified has 500 instances; the larger Test split provides higher statistical resolution.

Main findings.

  • The leading two Verified entries both resolve 396/500 instances; the top ten share 285 successes and 51 failures, leaving only 164 differentiating instances.
  • Median frontier solution-set nesting is 0.935 versus a score-implied baseline of 0.774.
  • Within-model scaffold ranges reach 29.8 percentage points, compared with an 8.8-point spread among the top thirty.
  • Exact paired McNemar tests distinguish none of the 29 adjacent top-thirty pairs on Verified at α=0.05; the larger Test split distinguishes 14 of 23.

Novelty vs. prior work.
The contribution is not a new agent but an evaluation audit that quantifies leaderboard resolution and scaffold dependence using paired outcomes. It turns a familiar qualitative concern—“small leaderboard gaps may be noise”—into a concrete protocol and empirical diagnosis.

Why it matters for AI4SE.
SWE-bench has become a de facto progress metric for coding agents. If its top entries cannot statistically support strict ordering, papers should avoid SOTA claims based on tiny score gaps and report model–scaffold provenance, paired comparisons, confidence/resolution, and larger or complementary evaluation sets.

Limitations / concerns.

  • The design is observational; model–scaffold interaction estimates do not establish causal scaffold effects.
  • Public submissions may be selected, tuned, or incompletely documented, so leaderboard provenance itself can bias the audit.
  • Non-rejection in McNemar tests is not evidence of equivalence; the paper explicitly cautions against that interpretation.

Reading recommendation.
Read now — especially the paired-test methodology, model–scaffold analysis, and five-step audit protocol; these should influence how future SWE-bench results are reported.


[P1] RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

  • Work ID: arxiv-2609.16936
  • Registry: record
  • Links: Paper · PDF source
  • Primary area: repository navigation / coding agents
  • Tags: multimodal code-graph context-management
  • Status: Preprint — venue not stated
  • Venue: arXiv
  • Version / date: arXiv v1, 2026-09-15
  • Authors: Yunxiang Zhang, Haiquan Wang, JiaWei Guo, Hanyang Xia, Yan Chen, Tong Chen, Zhang Zhiwei, Junchen Ye
  • Affiliations / team: not asserted in registry
  • Notable author/team signal: —
  • First seen: 2026-09-18

Fast grasp

One-sentence takeaway.
RepoAtlas is a training-free context module that repeatedly selects a task-relevant code-graph subregion and projects it into visual plus textual views, improving SWE-bench Verified resolve rate by 2.4 points while reducing input tokens by 5.8% and model calls by 7.8% versus the strongest multimodal graph baseline.

Research problem.
How can coding agents preserve non-local repository structure without dumping a dense full graph into context or relying on a local view that becomes stale as exploration proceeds?

Problem definition / setting.
During repository-level issue resolution, the agent explores files and symbols over time. A supporting module must choose a bounded representation of the relevant repository region and update it as the agent's evidence/state changes.

Core idea.
Maintain an evolving multimodal repository map through a select–project–refresh loop: issue evidence and exploration state choose a code-graph region, which is rendered into complementary visual/textual representations and refreshed when stale.

Method / system.

  • Builds/uses a repository code graph capturing non-local structure.
  • Selects a task-relevant subgraph under a fixed context budget using issue and exploration evidence.
  • Projects the subgraph into both visual and textual views for a multimodal coding agent.
  • Refreshes the view as exploration changes rather than treating retrieval as one-shot.

Evaluation.

  • Benchmarks / datasets: SWE-bench Verified.
  • Baselines: strongest reported multimodal graph baseline plus comparisons across three VLM families/scales.
  • Metrics: resolve rate, input-token use, model-call count.
  • Scale / setup: three multimodal models of different families and scales.

Main findings.

  • Resolve rate improves by 2.4 percentage points relative to the strongest multimodal graph baseline.
  • Average input tokens fall by 5.8% and model calls by 7.8%.
  • Gains are reported consistently across three models, suggesting the module is not tied to one VLM family.

Novelty vs. prior work.
The distinctive idea is not simply graph retrieval or visualization, but a state-dependent evolving view that combines graph topology with the agent's current exploration trajectory.

Why it matters for AI4SE.
Repository navigation and context selection remain major bottlenecks in coding agents. RepoAtlas suggests that multimodality can be useful not for screenshots alone but for compact structural representations of codebases, with efficiency gains as well as task gains.

Limitations / concerns.

  • Evaluation is limited to SWE-bench Verified; transfer to larger multilingual or long-horizon maintenance tasks is unclear.
  • A 2.4-point gain is modest and should be interpreted with the statistical-resolution concerns highlighted by today's SWE-bench audit.
  • The benefit depends on VLM competence at reading graph visualizations and on the quality of the underlying code graph.

Reading recommendation.
Add to queue — focus on graph selection/refresh criteria and ablations separating visual, textual, and refresh components.


3. Other new selected papers

No P2 papers selected today.


4. Publication and version updates for known works

No meaningful publication/version updates were confirmed for already indexed works in today's search window.


5. Possible duplicates requiring review

None.


6. Reading-queue actions

Add

  • arxiv-2609.18805 — ProgramDistill — new reference-guided, replay-verifiable coding-agent benchmark — priority P0
  • arxiv-2609.17394 — Coding Agents Have Converged — directly changes interpretation/reporting of SWE-bench leaderboard results — priority P0
  • arxiv-2609.16936 — RepoAtlas — adaptive multimodal repository context for coding agents — priority P1

Promote to deep read

  • arxiv-2609.18805 — benchmark construction and executable reference specification are unusually relevant to scalable agent evaluation/training.
  • arxiv-2609.17394 — statistical audit protocol should inform future benchmark reporting and SOTA claims.

Remove / deprioritize

  • None.

7. Coverage and search notes


Sources checked

Software Engineering journals / venues

  • [x] TSE
  • [x] TOSEM
  • [x] EMSE
  • [x] JSS
  • [x] ICSE
  • [x] FSE / ESEC-FSE
  • [x] ASE
  • [x] ISSTA
  • [x] MSR
  • [x] ACM SIGSOFT publications / relevant SIGSOFT venues

General AI / ML / NLP venues

  • [x] ICML
  • [x] NeurIPS
  • [x] ICLR
  • [x] ACL
  • [x] AAAI

arXiv

  • [x] cs.SE
  • [x] cs.AI intersecting software engineering
  • [x] cs.CL intersecting software engineering
  • [x] cs.LG intersecting software engineering

High-value signals

  • [x] Notable researchers / labs
  • [x] Major technology companies
  • [x] New benchmark or dataset releases
  • [x] New code/model releases attached to recent papers

Coverage gaps / failures: no blocking source failure. Publisher/venue status was kept conservative; ADMA acceptance is recorded only because the arXiv record explicitly states it.


8. Notes

Today's three papers fit together unusually well: ProgramDistill broadens what constitutes a specification, RepoAtlas changes how an agent sees repository state, and the SWE-bench audit changes how we should interpret the resulting score. For AI4SE experiments, benchmark validity, context/harness provenance, and statistical resolution increasingly need to be treated as first-class methodological variables rather than afterthoughts.