Daily Paper Digest — 2026-09-19
目录
Discovery window: 2026-09-16 – 2026-09-19
Registry snapshot: schema v2, 19 known works
Selected new works: 3 · P0: 2 · P1: 1
Today's signal
Today's strongest signal is harness engineering becoming an independent research variable rather than an implementation detail. One paper isolates planning, action space, and context management inside coding-agent harnesses; another turns scientific repositories into executable learning environments; a third shows that a safety-aware harness can change whether generated robot-control code respects physical constraints. Across all three, the model alone is not the unit of capability: environment design, context policy, verification, and constraint enforcement materially determine observed performance.
Must-read shortlist
| Priority | Paper | Area | Status | Why it matters |
|---|---|---|---|---|
| P0 | An Empirical Study of Harness Design for Coding Agents | coding agents / harness engineering | Preprint | 176 matched settings isolate context management, planning and action-space effects on SWE-Bench Verified and Terminal-Bench 2.1. |
| P0 | ScienceIDE | agent training / scientific software | Preprint | Converts real scientific repositories into programmable, verifiable SFT/RL/evaluation environments and trains 4B/9B/72B models from verified trajectories. |
| P1 | Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation | coding agents / safety / embodied code | Preprint | Shows explicit safety instructions are insufficient unless the harness operationalizes them in planning and contact execution. |
P0/P1 priority papers
[P0] An Empirical Study of Harness Design for Coding Agents
- Work ID:
arxiv-2609.20804 - Version: arXiv v1, 2026-09-17
- Authors: Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang
- Status / venue: Preprint — venue not stated
- Area: coding agents, harness engineering, context management
- Code/project/data: no authoritative artifact link asserted from the arXiv landing page
Research question. Which coding-agent harness components actually contribute to long-horizon software-engineering performance, and how do those contributions depend on model capability and context budget?
Problem definition / setting. Prior comparisons usually bundle scaffold choices together. This study fixes a lightweight execution loop and varies three dimensions independently: planning, action space, and context management.
Core idea and method. Evaluate matched harness configurations rather than monolithic agent products. The study covers 176 matched settings across four models, five context-management strategies, four context-window budgets, plus targeted planning/action-space ablations.
Benchmarks / baselines / metrics. SWE-Bench Verified and Terminal-Bench 2.1; comparisons are within a fixed harness under component substitutions, with task accuracy, failure behavior and cost/efficiency as central outcomes.
Main results. Context management matters increasingly as the context budget tightens, primarily by preventing overflow. Rule-based elision followed by LLM summarization gives the best overall efficiency; recoverable elision adds machinery models rarely use and gives no accuracy gain. Planning changes role with model strength: it scaffolds weaker models but mostly saves cost for stronger models. Predefined tools help models with weak bash proficiency, while bash-capable models can use a bash-only action space more cheaply, especially on command-line-heavy tasks.
Real novelty. The paper changes the experimental unit from “agent framework” to individual harness mechanisms, producing model- and budget-conditional design guidance rather than another leaderboard point.
Why it matters for AI4SE. It directly affects how SWE-bench results should be interpreted and how coding agents should be engineered: model quality and harness quality are coupled, and a component that helps one model/budget can be redundant for another.
Limitations. Results are tied to four models, two benchmarks and one fixed lightweight loop. Component interactions outside the tested factorial slices may differ in richer production agents; cost conclusions also depend on model/tool pricing and tokenization.
Reading recommendation. Read now. Prioritize the matched-setting methodology, context-overflow analysis, and model-strength interaction plots; these are reusable for coding-agent experimental design.
[P0] ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
- Work ID:
arxiv-2609.19134 - Version: arXiv v1, 2026-09-16
- Authors: Hejia Geng et al. (45 authors)
- Status / venue: Preprint — venue not stated
- Area: scientific software, agent environments, execution-grounded training
- Code: https://github.com/AI-Frontier-Labs/ScienceIDE
Research question. Can real scientific software be systematically converted from difficult-to-use repositories into verifiable learning environments for agents, rather than merely serving as static code corpora?
Problem definition / setting. Scientific repositories contain executable domain knowledge but also fragmented toolchains, implicit conventions and specialized correctness criteria. These properties create a “scientific experience bottleneck” for SFT, RL and evaluation.
Core idea and method. Experts define scientific cases and acceptance criteria; agents then transform repositories into executable environments supporting task generation, execution and scientific verification. Verified interaction trajectories become training data for PhAI-IDE-4B/9B/72B.
Benchmarks / datasets / metrics. The central evaluation is held-out scientific-code repair, supplemented by selected general-purpose code, reasoning and knowledge benchmarks to test transfer. Verification is environment/execution based rather than diff matching.
Main results. The PhAI-IDE model family improves on held-out scientific-code repair and shows positive transfer on selected general-purpose benchmarks, supporting the claim that executable scientific experience can train capabilities beyond the source domains.
Real novelty. The contribution is infrastructure: scientific repositories become programmable agent environments with domain-aware acceptance criteria, connecting repository mining, task generation, verification and training in one loop.
Why it matters for AI4SE. This is a direct bridge between software repositories and verifiable agent training. It complements SWE-bench Science: instead of only measuring repair capability, it asks how repositories can supply scalable training experience with stronger semantic judges.
Limitations. Environment conversion depends on expert-defined cases/criteria and on repositories that can be made executable. Positive transfer is benchmark-dependent, and the 45-author project spans heterogeneous scientific domains, making attribution of gains to specific environment properties difficult.
Reading recommendation. Read now. Focus on environment construction, acceptance criteria, task generation, and the evidence connecting verified trajectories to transfer.
[P1] Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation
- Work ID:
arxiv-2609.20822 - Version: arXiv v1, 2026-09-17
- Authors: Bingxin Xu, Yuzhang Shang, Zhen Dong, Emilio Ferrara
- Status / venue: Preprint — venue not stated
- Area: coding agents, safety, embodied software
Research question. When a coding agent writes robot controllers, does a natural-language safety constraint actually influence planning and execution, and can harness design enforce it?
Setting and idea. Each task pairs a manipulation goal with an obstacle that must not be touched. Baseline agents often reason about the obstacle yet still collide because the constraint is not promoted into planning priority. SafeHarness separates route planning from contact-rich execution and makes obstacle avoidance explicit in both.
Method / evaluation. The route harness grounds objects as bounding boxes, proposes waypoint routes, verifies them and replans when needed; contact execution chooses a safe contact position. Reported outcomes are task success and collision avoidance against the same agent without these harnesses and prior SOTA.
Main results. SafeHarness reaches 71.9% task success and 87.5% collision avoidance, improving over prior SOTA by 6.5 and 27.0 percentage points respectively; these are 2.3× and 1.5× the same agent without harnesses.
Novelty / AI4SE relevance. Although robotics-first, it is a useful SE-adjacent example of constraints being ineffective when they remain prompt text. For coding agents operating deployment, CI/CD or infrastructure tools, safety policies may likewise need executable harness enforcement rather than instruction-only alignment.
Limitations. The setting is embodied manipulation, not conventional software repositories, and obstacle geometry gives a relatively concrete safety specification. Transfer to security, deployment or code-change constraints remains a hypothesis.
Reading recommendation. Queue. Read primarily for the constraint-to-harness design pattern, not as a core SE benchmark.
Other new papers
No additional paper cleared today's P0/P1 inclusion threshold after deduplication and relevance filtering.
Known-paper publication/version updates
No meaningful publication or version update for an indexed work was confirmed today.
Possible duplicates
None requiring manual confirmation.
Reading-queue actions
- Add
arxiv-2609.20804— P0, deep-read candidate. - Add
arxiv-2609.19134— P0, deep-read candidate. - Add
arxiv-2609.20822— P1, targeted read on safety harness design.
Coverage notes
Checked recent arXiv cs.SE and SE-relevant cs.AI/cs.CL/cs.LG signals, with targeted searches around coding agents, repository-level coding, SWE-bench, testing, code generation, agent evaluation and safety. Also screened for recent venue/publication signals across the requested SE/AI venue set. Publication status was kept conservative: all three selected works are recorded only as preprints because no authoritative venue acceptance/publication statement was confirmed.
评论 (0)