Discovery window: 2026-09-19 – 2026-09-23
Candidates checked: 12 · New works: 1 · Known unchanged: 3 · Version/publication updates: 0
Selected new works: 1 · P0: 0 · P1: 1 · P2: 0
Focus: AI4SE first; general AI only when it can materially affect software-engineering research or practice.

0. Deduplication summary

  • New works: 1
  • Suppressed as already known and unchanged: 3
  • Known works with meaningful updates: 0
  • Possible duplicates requiring identity check: 0

The first-pass deduplication database was papers/registry/index.json (schema_version=2, 27 works). Exact arXiv-ID and normalized-title checks suppressed already tracked works before analysis.


1. Today's signal

What is worth noticing today?

  • Deterministic runtime verification is moving beyond ordinary repositories. CraftBench-UE evaluates coding agents inside Unreal Engine and reconstructs saved submissions in fresh projects, making build, asset, and runtime behavior first-class correctness criteria rather than relying on an LLM judge.
  • Artifact validity is not behavioral correctness. A substantial fraction of Blueprint submissions that satisfy asset checks still fail explicit runtime assertions, echoing the recent SWE-Gate/GameLogicBench signal that surface-level success criteria overestimate coding-agent capability.
  • Representation and tooling matter. On paired tasks with identical gameplay requirements and runtime tests, C++ completion materially exceeds Blueprint completion, suggesting current agents remain sensitive to authoring representation and editor affordances.

Must-read shortlist

PriorityPaperAreaStatus / VenueWhy it mattersAction
P1CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Enginecoding-agents benchmark runtime-verificationPreprint — venue not statedExtends deterministic coding-agent evaluation to mixed code/assets/editor state in a real game engine.Queue

2. New priority papers

[P1] CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

  • Work ID: arxiv-2609.23142
  • Registry: record
  • Links: Paper · PDF source
  • Primary area: coding-agent evaluation
  • Tags: coding-agents benchmark Unreal-Engine runtime-verification
  • Status: Preprint — venue not stated
  • Venue: arXiv
  • Version / date: arXiv v1, 2026-09-19
  • Authors: Shutong Wu, Kevin Calderone, Andy Tsen
  • Affiliations / team: not established from the authoritative metadata checked today
  • Notable author/team signal: —
  • First seen: 2026-09-23

Fast grasp

One-sentence takeaway.
CraftBench-UE introduces a deterministic Unreal Engine evaluation harness and 70-task benchmark showing that coding agents can produce structurally valid engine artifacts that still fail the requested runtime behavior, with a pronounced C++/Blueprint performance gap.

Research problem.
How should coding agents be evaluated when software changes include source code, visual assets, editor state, and runtime gameplay behavior, where compilation or an LLM judge is insufficient?

Problem definition / setting.
Agents operate in an isolated Unreal Engine environment and must complete gameplay-development tasks spanning C++ source, Blueprint assets, and editor scripting. Submitted projects are reconstructed in fresh projects and evaluated deterministically for build validity, asset requirements, and runtime behavior.

Core idea.
Treat the game engine itself as a verifiable software-engineering environment: separate artifact/build checks from explicit runtime assertions, and pair equivalent C++ and Blueprint tasks to isolate the effect of implementation representation and tool access.

Method / system.

  • Isolated Unreal Engine agent environment with saved-submission reconstruction in fresh projects.
  • Deterministic build, asset, and runtime checks; no LLM judge in the evaluation loop.
  • 70 tasks covering C++, Blueprint, and editor scripting.
  • Seven models evaluated under two editor-tool configurations, plus a file-and-shell baseline for C++ tasks.
  • Ten paired tasks require the same gameplay and use the same runtime tests while changing the required deliverable between C++ and Blueprint.

Evaluation.

  • Benchmarks / datasets: CraftBench-UE, 70 Unreal Engine tasks.
  • Baselines: seven evaluated models; two editor-tool configurations; file-and-shell baseline on C++ tasks.
  • Metrics: task completion plus deterministic build/asset/runtime checks.
  • Scale / setup: 70 tasks; 10 representation-paired tasks.

Main findings.

  • Across the 10 paired tasks, C++ completion exceeds Blueprint completion by 30.0 and 42.9 percentage points under the two tool configurations.
  • Among on-time Blueprint submissions in the paired set that pass asset checks, 42.2% and 50.0% still fail explicit runtime assertions.
  • The result demonstrates a concrete evaluator gap: structurally valid engine artifacts can remain behaviorally wrong.

Novelty vs. prior work.
The main novelty is a deterministic, engine-native benchmark that jointly evaluates source code, visual/editor artifacts, and runtime gameplay behavior, together with controlled C++/Blueprint paired tasks. This extends repository-centric coding-agent benchmarks into stateful IDE/game-engine development without falling back to subjective LLM judging.

Why it matters for AI4SE.
The work is useful for coding-agent benchmark design, test-oracle construction, multimodal/editor tool use, and evaluation of non-textual software artifacts. Its strongest connection to recent AI4SE work is the distinction between superficial artifact validity and end-to-end behavioral correctness.

Limitations / concerns.

  • The benchmark is specific to Unreal Engine and contains 70 tasks; generalization to other complex IDEs and engines remains open.
  • The abstract reports aggregate representation gaps but does not by itself establish whether failures come primarily from model reasoning, Blueprint interaction ergonomics, tool APIs, or evaluator/task difficulty.
  • Code/harness release is announced with the report, but no separate authoritative project/code URL was established in today's check.

Reading recommendation.
Add to queue — prioritize the harness/evaluator construction, paired C++/Blueprint experimental design, and failure analysis; these are more reusable for AI4SE than the absolute leaderboard numbers.


3. Other new selected papers

No additional P2 papers met today's retention threshold.


4. Publication and version updates for known works

No meaningful publication/version updates were verified today.


5. Possible duplicates requiring review

None.


6. Reading-queue actions

Add

  • arxiv-2609.23142 — CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine — deterministic runtime evaluation for mixed code/editor artifacts — priority P1

Promote to deep read

  • None today.

Remove / deprioritize

  • None today.

7. Coverage and search notes


Sources checked

Software Engineering journals / venues

  • [x] TSE / TOSEM / EMSE / JSS recent-paper web signals
  • [x] ICSE / FSE / ASE / ISSTA / MSR / SIGSOFT recent-paper web signals

General AI / ML / NLP venues

  • [x] ICML / NeurIPS / ICLR / ACL / AAAI recent-paper web signals

arXiv

  • [x] cs.SE
  • [x] cs.AI intersecting software engineering
  • [x] cs.CL intersecting software engineering
  • [x] cs.LG intersecting software engineering

High-value signals

  • [x] Coding-agent benchmarks and deterministic evaluation
  • [x] Repository-level coding / testing / runtime verification
  • [x] Recent major-lab and company signals

Coverage gaps / failures: Search indexing for same-day arXiv submissions was incomplete; the scan therefore used a multi-day recent window and conservative selection. No venue status was inferred from timing or formatting.


8. Notes

Today's strongest cross-paper theme is evaluator design: recent benchmark work increasingly treats build success, asset validity, tests, and runtime semantics as separate layers rather than a single pass/fail oracle. CraftBench-UE is worth tracking as a concrete example outside conventional repository-only development.