Daily Paper Digest — 2026-09-15
目录
Discovery window: 2026-08-20 – 2026-09-15
Candidates checked: 13 · New works: 5 · Known unchanged: 8 · Version/publication updates: 0
Selected new works: 5 · P0: 2 · P1: 3 · P2: 0
Focus: AI4SE first; general AI only when it can materially affect software-engineering research or practice.
0. Deduplication summary
- New works: 5
- Suppressed as already known and unchanged: 8
- Known works with meaningful updates: 0
- Possible duplicates requiring identity check: 0
The first-pass deduplication used papers/registry/index.json (schema v2). Exact arXiv identifiers were checked first; no new candidate matched an existing work. Existing indexed works were suppressed rather than summarized again.
1. Today's signal
What is worth noticing today?
- Coding-agent evaluation is broadening beyond issue fixing. SWE-bench Science tests domain-heavy scientific repositories, while SWE Refactor Bench targets long-horizon whole-repository migrations and explicitly checks whether the requested transformation actually occurred.
- Execution feedback is becoming a first-class research object. AMDKernelVault uses compilation, execution and latency feedback for code-model training, while ParaRecover evaluates whether agents can localize and recover from intermediate tool-use failures rather than merely finish the task.
- “Tests passed” is increasingly treated as insufficient evidence. The new testing survey reinforces the same theme seen in recent agent benchmarks: test availability, validity, feedback use and evaluation independence need to be separated rather than collapsed into one pass/fail signal.
Must-read shortlist
| Priority | Paper | Area | Status / Venue | Why it matters | Action |
|---|---|---|---|---|---|
| P0 | SWE-bench Science | coding-agents benchmark | Preprint — venue not stated | Extends repository-level agent evaluation into scientific software, where domain contracts and scientific knowledge materially affect repair. | Read / Queue |
| P0 | SWE Refactor Bench | coding-agents maintenance | Preprint — venue not stated | Evaluates whole-repository migrations with a three-stage verifier that distinguishes real migration from behavioural test gaming. | Read / Queue |
| P1 | AMDKernelVault | code-generation agentic-training | Preprint — venue not stated | Large open AMD kernel corpus plus execution-aware SFT/RL for code optimization beyond CUDA-centric settings. | Queue |
| P1 | ParaRecover | agents evaluation | Preprint — venue not stated | Measures process-level failure localization and recovery in parallel tool-use trajectories. | Queue |
| P1 | Test-Driven Approaches to Software Engineering with Large Language Models | testing survey | Preprint — venue not stated | Useful synthesis of how tests condition generation, repair, analysis and agent skills; clarifies common evaluation confounds. | Queue |
2. New priority papers
[P0] SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
- Work ID:
arxiv-2608.19799 - Registry: record
- Links: Paper · PDF source · Code · Project · Data
- Primary area:
coding agents / repository-level software engineering - Tags:
SWE-benchscientific-softwarebenchmark - Status: Preprint — venue not stated
- Venue: arXiv
- Version / date: arXiv v1, 2026-08-20
- Authors: Zhipeng Xu, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu
- Affiliations / team: Shanghai Innovation Institute / Fudan University; OpenMOSS project
- Notable author/team signal: Xipeng Qiu's group has substantial LLM and agent research visibility; the benchmark has an open code/data/leaderboard stack.
- First seen: 2026-09-15
Fast grasp
One-sentence takeaway.
A repository-level benchmark of 119 tasks from 98 scientific repositories shows that even strong coding agents remain below 50% pass@1 and fail in ways tied to scientific abstractions, exploration quality, repair coverage and domain-knowledge use.
Research problem.
Do current coding agents reliably repair scientific software, where correctness depends not only on ordinary program behavior but also on domain-specific contracts such as units, numerical invariants, data formats and scientific assumptions?
Problem definition / setting.
The benchmark spans 20 scientific domains and organizes tasks into issue-driven, expert-exploratory and engineering-integration paradigms. Agents must modify real repositories while respecting both executable engineering constraints and domain-specific scientific semantics.
Core idea.
The paper moves agent evaluation from generic repository repair toward domain-grounded software engineering, and pairs aggregate success with a failure taxonomy plus an ablation on scientific guidance.
Method / system.
- Curates 119 executable tasks from 98 scientific repositories across 20 domains.
- Defines three task paradigms to cover explicit issue repair, exploratory engineering and integration work.
- Evaluates coding agents under repository-level execution.
- Performs failure analysis and a paired ablation removing explicit scientific guidance while holding the engineering context fixed.
Evaluation.
- Benchmarks / datasets: SWE-bench Science, 119 tasks / 98 repositories / 20 domains.
- Baselines: multiple coding agents; the reported best configuration is Claude Code with Opus-5 (max).
- Metrics: primarily pass@1, plus failure-mechanism analysis and token-efficiency observations.
- Scale / setup: real GitHub scientific repositories with executable task environments.
Main findings.
- Best reported agent performance remains below 50% pass@1.
- Four recurring failure modes dominate: missing scientific knowledge/abstraction, misguided exploration or shallow repair, incomplete repair/integration, and failure to generalize scientific knowledge.
- Scientific guidance helps only when well grounded; poorly aligned guidance can anchor the agent and does not reliably improve exact repair success.
Novelty vs. prior work.
The strongest novelty is the task setting: scientific repositories expose semantic contracts that ordinary SWE-bench-style tasks often do not. The paired guidance ablation also gives a concrete empirical handle on when domain knowledge helps versus harms.
Why it matters for AI4SE.
This is directly relevant to repository-level issue resolution, localization and repair. It suggests that future coding-agent systems need explicit mechanisms for domain knowledge acquisition, validation and uncertainty handling rather than assuming generic code reasoning is sufficient.
Limitations / concerns.
- Scientific repositories are highly heterogeneous; task construction and domain coverage may still favor projects that can be containerized and automatically evaluated.
- The benchmark is small relative to general SWE-bench derivatives and may not support fine-grained conclusions for each scientific domain.
- Domain guidance quality is itself a confound: the result that “knowledge can hurt” depends strongly on how guidance is selected and presented.
Reading recommendation.
Read now — focus on benchmark construction, the four failure categories and the scientific-guidance ablation; these are the most reusable parts for coding-agent evaluation research.
[P0] SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
- Work ID:
arxiv-2608.23564 - Registry: record
- Links: Paper · PDF source · Code · Project
- Primary area:
coding agents / maintenance / migration - Tags:
repository-levelrefactoringbenchmark - Status: Preprint — venue not stated
- Venue: arXiv
- Version / date: arXiv v1, 2026-08-24
- Authors: Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
- Affiliations / team: not asserted in the registry; project released by Einsia
- Notable author/team signal: strong benchmark/artifact signal: 520 graded runs, released score tables, trajectories and task images.
- First seen: 2026-09-15
Fast grasp
One-sentence takeaway.
A 20-task benchmark for whole-repository migrations finds that only 5.4% of 520 agent runs pass migration audit, behavioral tests and adversarial agentic verification together.
Research problem.
Can coding agents perform long-horizon repository migrations rather than merely produce patches that preserve tests while sidestepping the requested architectural or technology change?
Problem definition / setting.
Each task asks an agent to migrate an entire repository across one of four technical-debt categories. Success requires both completing the intended migration and preserving behavior.
Core idea.
The key contribution is an evaluation protocol that separates migration completeness from behavioral correctness, closing a benchmark loophole where an agent can pass tests by retaining or copying the old implementation.
Method / system.
- Builds 20 whole-repository migration tasks across four technical-debt categories.
- Stage 1, Migration Audit, verifies that the requested migration actually occurred.
- Stage 2, Behavioral Tests, evaluates correctness on a fixed suite.
- Stage 3, Agentic Verification, uses six independent coding agents to generate targeted tests for hidden behavioral differences.
Evaluation.
- Benchmarks / datasets: 20 migration tasks.
- Baselines: 8 frontier models and 26 model-effort configurations.
- Metrics: three-stage pass result and composite score; task/category breakdowns.
- Scale / setup: 520 total runs; released trajectories include roughly 248k tool calls according to the project site.
Main findings.
- Only 28/520 runs (5.4%) pass all three stages; 13/20 tasks have no accepted solution.
- The best model scores 47.0/100.
- Of 340 runs that pass the migration audit, 58% reach 99% of fixed checks but only 26% reach 100%, showing a severe “last mile” reliability problem.
- Build-toolchain rewrites are substantially easier than language rewrites (reported category scores 31.4 vs. 5.6).
Novelty vs. prior work.
Rather than another issue-fixing benchmark, this work targets long-horizon maintenance and explicitly verifies that the requested transformation happened. The agentic verifier is also an interesting attempt to find behavioral differences beyond a frozen test suite.
Why it matters for AI4SE.
It exposes a benchmark-design problem that also appears in APR and issue resolution: behavioral tests can be satisfied by solutions that violate the intended change. This is highly relevant to trustworthy coding-agent evaluation and specification-aware repair.
Limitations / concerns.
- Only 20 tasks, so category-level generalization should be treated cautiously.
- Agentic verification depends on the quality and diversity of the six verifier agents and may itself miss hidden behavior.
- Whole-repository migrations are expensive, making broad model/harness replication costly.
Reading recommendation.
Read now — prioritize the three-stage verifier design, examples of “blindness,” and the error distribution after Migration Audit; these are directly useful for benchmark methodology.
[P1] AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization
- Work ID:
arxiv-2609.12471 - Registry: record
- Links: Paper · PDF source · Code · Data
- Primary area:
code generation / code foundation models - Tags:
GPU-kernelsexecution-feedbackRL - Status: Preprint — venue not stated
- Venue: arXiv
- Version / date: arXiv v1, 2026-09-11
- Authors: Ji Liu, Saptarshi Majumder, Yiqing Huang, Wenwen Ouyang, Umang Pandey, Zeping Li, Chushi Chen, Zihao An, Puyuan Yang, Zekai Li, Sina Rafati, Ziqiong Liu, Pratik Prabhanjan Brahma, Dong Li, Zicheng Liu, Sharon Zhou, Emad Barsoum
- Affiliations / team: AMD
- Notable author/team signal: direct industrial hardware/compiler stack ownership plus a fully open dataset/code release.
- First seen: 2026-09-15
Fast grasp
One-sentence takeaway.
AMD releases a large execution-verified HIP/Triton corpus and agentic generation pipeline, then shows that an 8B model trained with SFT plus execution-aware RL can lead compared models on several AMD-kernel correctness benchmarks under fixed budgets.
Research problem.
Most LLM kernel-generation work is CUDA/NVIDIA-centric and often depends on repeated calls to frontier models; the paper asks whether open data and execution-aware training can support capable AMD-targeted kernel generation.
Problem definition / setting.
Given PyTorch references, generate HIP or Triton GPU kernels for recent AMD CDNA hardware, compile and validate them under ROCm, and optimize latency while preserving correctness.
Core idea.
Use agent-driven synthesis pipelines to create a large, execution-verified training corpus, then distill this process into a smaller model with supervised and execution-aware reinforcement learning.
Method / system.
- HIPKernelGen and TritonKernelGen synthesize kernels from PyTorch references.
- Candidate kernels are compiled and execution-validated under ROCm.
- Valid kernels are latency-profiled on AMD hardware.
- Qwen3-8B is trained with SFT followed by execution-aware RL.
Evaluation.
- Datasets: 62,153 verified HIP samples, 2,377 production-grounded ROCm Libraries QA entries, and 39,893 Triton kernels.
- Benchmarks: PyTorch-to-HIP, TritonBench-G, ROCmBench.
- Metrics: Pass@1 / Corr@3 plus compilation and speed metrics.
- Scale / setup: fixed evaluation budgets on AMD CDNA/ROCm settings.
Main findings.
- 34.0% Pass@1 on PyTorch-to-HIP.
- 33.2% Corr@3 on TritonBench-G.
- 41.94% Corr@3 on ROCmBench.
- The trained model does not uniformly dominate compilation success or performance, so correctness gains do not imply universal optimization gains.
Novelty vs. prior work.
The main novelty is ecosystem coverage and artifact scale: a large AMD-native corpus and end-to-end agentic data-generation/training pipeline, rather than another CUDA-only kernel benchmark.
Why it matters for AI4SE.
This is a concrete example of execution-grounded code-model training where compilers, tests and performance measurements are all part of the feedback loop. The pattern is transferable to repository-level code generation and repair.
Limitations / concerns.
- Highly specialized domain; success may not transfer to general software engineering.
- Kernel correctness and latency depend strongly on hardware/compiler versions.
- The model does not consistently win on speed metrics, so optimization quality remains distinct from functional correctness.
Reading recommendation.
Add to queue — especially useful for execution-aware training pipelines, code-data construction and compiler-in-the-loop agent design.
[P1] ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents
- Work ID:
arxiv-2609.12345 - Registry: record
- Links: Paper · PDF source · Code/Data
- Primary area:
agent evaluation / reliability - Tags:
tool-useerror-localizationreplanning - Status: Preprint — venue not stated
- Venue: arXiv
- Version / date: arXiv v1, 2026-09-11
- Authors: Bowen Guan, Zhentao Yin, Yanming Shen
- Affiliations / team: Dalian University of Technology
- Notable author/team signal: —
- First seen: 2026-09-15
Fast grasp
One-sentence takeaway.
ParaRecover shows that high end-task pass rates can coexist with weak process reliability: models often fail to correctly diagnose dependencies, implicit failures and recovery strategies in parallel multi-turn tool execution.
Research problem.
Can tool-using agents detect where an intermediate execution failed, reason about downstream effects and replan correctly when failures propagate across parallel branches?
Problem definition / setting.
Agent execution is represented as a DAG of dependent subtasks/tool calls. The agent receives a partially executed graph plus observations and must localize failures, infer the executable frontier and produce a corrected execution plan.
Core idea.
Evaluate recovery inside the trajectory rather than scoring only final success, using a fine-grained error taxonomy and the SDE rubric: Structural Integrity, Diagnostic Reasoning and Evolutionary Strategy.
Method / system.
- Defines 14 error types over planning dependencies, tool selection and argument matching.
- Builds 10,626 instances at two difficulty levels.
- LEVEL-1 injects recent-step errors; LEVEL-2 includes multi-turn propagation.
- Scores both process reasoning and final pass behavior.
Evaluation.
- Benchmarks / datasets: 10,626 ParaRecover instances.
- Baselines: more than ten mainstream model families, including OpenAI, Anthropic, Google, Qwen, GLM and DeepSeek variants.
- Metrics: SI, DR, ES, average SDE score and Pass@1.
- Scale / setup: parallel DAG-based multi-turn tool-use trajectories.
Main findings.
- Models can exceed 90% Pass@1 while having much lower SDE process scores.
- Multi-turn propagation particularly hurts structural/dependency reasoning and recovery strategy.
- The paper reports that SDE-derived supervision can improve reflective recovery capability.
Novelty vs. prior work.
The key novelty is process-level recovery evaluation for parallel tool use, including explicit error localization and corrective replanning rather than only tool-call correctness or final success.
Why it matters for AI4SE.
Coding agents routinely run tests, linters, search tools and shell commands in branching workflows. A process metric like this could complement SWE-bench-style final patch success and expose brittle recovery behavior.
Limitations / concerns.
- The benchmark is general-agent rather than software-engineering-specific.
- DAG abstractions simplify messy real coding trajectories, where dependencies are often implicit and stateful.
- Some SDE dimensions rely on rubric-style judgment and may inherit evaluator-model biases.
Reading recommendation.
Add to queue — useful for designing trajectory-aware evaluation or fault-localization mechanisms inside coding agents.
[P1] Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills
- Work ID:
arxiv-2609.12012 - Registry: record
- Links: Paper · PDF source
- Primary area:
software testing / LLM4SE - Tags:
test-drivenagentssurvey - Status: Preprint — venue not stated
- Venue: arXiv
- Version / date: arXiv v1, 2026-09-10
- Authors: Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni
- Affiliations / team: not asserted
- Notable author/team signal: —
- First seen: 2026-09-15
Fast grasp
One-sentence takeaway.
The survey reframes “test-driven LLM software engineering” around the decision changed by tests, separating true Red-Green-Refactor behavior from test-conditioned generation, execution-guided refinement, test-mediated analysis and evaluation-only testing.
Research problem.
The literature often labels many different uses of tests as test-driven or execution-guided; the paper asks how to distinguish these mechanisms and what evidence is actually available across LLM-based SE tasks.
Problem definition / setting.
A structured scoping review of 87 research/supporting records, with method/protocol-level extraction for 83 records plus five practice resources, spanning generation, repair, translation, refactoring, clone detection, search, localization, training-data construction and formal-specification validation.
Core idea.
Organize work by what decision the test changes and treat test availability, validity, feedback use and evaluation independence as separate properties.
Method / system.
- Separates canonical Red-Green-Refactor from adjacent test-using workflows.
- Builds a mechanism taxonomy across multiple SE tasks.
- Includes an explicit analysis of reusable agent skills that encode testing procedures.
- Examines protocol-level evidence rather than relying only on paper-level labels.
Evaluation.
- Corpus: 87 research/supporting records; 83 receive detailed method/protocol extraction.
- Practice resources: 5.
- Tasks covered: code generation, repair, translation, refactoring, clone detection, search, localization, training-data construction and formal validation.
- Metrics: not a model benchmark; evidence is synthesized across reported protocols and outcomes.
Main findings.
- Test passing alone does not demonstrate behavioral equivalence, effective feedback use or adherence to a test-driven process.
- Aggregate improvements can conceal materially different behaviors across models, tasks and denominators.
- Oracle validation, causal evaluation, long-horizon maintenance and reusable test-driven agent capabilities remain open research directions.
Novelty vs. prior work.
Its useful contribution is conceptual hygiene: it distinguishes several mechanisms that are often collapsed into “test-driven” LLM engineering and emphasizes evaluation independence.
Why it matters for AI4SE.
This is directly relevant to APR, repository-level coding agents and test-generation research, especially when designing benchmarks where generated tests or visible tests can leak into the optimization loop.
Limitations / concerns.
- As a scoping survey, conclusions depend on categorization choices and heterogeneous primary-study reporting.
- Rapidly evolving agent tooling means the practice side can become stale quickly.
- It provides taxonomy and synthesis rather than causal evidence for which testing strategy works best.
Reading recommendation.
Add to queue — read the taxonomy and protocol/evaluation sections first; they are useful for designing or reviewing LLM4SE experiments.
3. Other new selected papers
No P2 papers selected today.
4. Publication and version updates for known works
No meaningful publication/version updates were confirmed for already indexed works in today's search window.
5. Possible duplicates requiring review
None.
6. Reading-queue actions
Add
arxiv-2608.19799— SWE-bench Science — directly relevant new repository-level coding-agent benchmark — priority P0arxiv-2608.23564— SWE Refactor Bench — strong long-horizon maintenance benchmark with migration-aware verification — priority P0arxiv-2609.12471— AMDKernelVault — execution-aware code generation/training pipeline and open data — priority P1arxiv-2609.12345— ParaRecover — trajectory-level agent recovery evaluation relevant to coding-agent reliability — priority P1arxiv-2609.12012— Test-Driven Approaches... — useful LLM4SE testing taxonomy and evaluation guidance — priority P1
Promote to deep read
arxiv-2608.19799— benchmark construction and scientific-guidance ablation are directly relevant to repository-level agent research.arxiv-2608.23564— three-stage verification design is highly relevant to benchmark validity and long-horizon agent evaluation.
Remove / deprioritize
- None.
7. Coverage and search notes
Sources checked
Software Engineering journals / venues
- [x] TSE
- [x] TOSEM
- [x] EMSE
- [x] JSS
- [x] ICSE
- [x] FSE / ESEC-FSE
- [x] ASE
- [x] ISSTA
- [x] MSR
- [x] ACM SIGSOFT publications / relevant SIGSOFT venues
General AI / ML / NLP venues
- [x] ICML
- [x] NeurIPS
- [x] ICLR
- [x] ACL
- [x] AAAI
arXiv
- [x] cs.SE
- [x] cs.AI intersecting software engineering
- [x] cs.CL intersecting software engineering
- [x] cs.LG intersecting software engineering
High-value signals
- [x] Notable researchers / labs
- [x] Major technology companies
- [x] New benchmark or dataset releases
- [x] New code/model releases attached to recent papers
Coverage gaps / failures: broad venue searches were completed, but same-day indexing lag means very recent publisher pages may appear later than arXiv/author pages. Publication status was therefore kept conservative unless explicitly stated.
8. Notes
A consistent cross-paper theme is that final behavioral success is becoming an inadequate single metric for agentic software engineering. Scientific-domain constraints, migration completeness, execution-process reliability and test independence all expose dimensions that ordinary “tests pass” scoring can miss.
评论 (0)