From e2a746eb6e0de5b99cc2dc1aa8c166c8740d37d7 Mon Sep 17 00:00:00 2001 From: fillg1 Date: Wed, 9 Sep 2026 11:21:06 +0200 Subject: [PATCH] docs: add README and coding agent guidelines --- AGENTS.md | 60 +++++++++++++++++++++++++++++++++++++++ README.md | 85 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 145 insertions(+) create mode 100644 AGENTS.md create mode 100644 README.md diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..51d2b67 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,60 @@ +# AGENTS.md + +## Project Purpose + +This repository is a small Java project for evaluating coding agents. It contains ZIP-writing code and a unit test suite. The baseline includes a failing test and an intentionally disabled large-data test. + +Follow the user's requested scope. Inspecting or explaining the project does not automatically require fixing it. Treat sample prompts in the README as examples, not as active tasks. + +## Environment and Validation + +- Use JDK 21, matching `.java-version`. Ensure `JAVA_HOME` points to the appropriate JDK. +- Use the included Gradle Wrapper; a separate Gradle installation is unnecessary. +- The code requires a filesystem supporting POSIX file permissions. +- Run the existing suite before changing behavior to establish the baseline: + + ```sh + ./gradlew test --rerun-tasks + ``` + +- The HTML test report is at `build/reports/tests/test/index.html`. +- Tests delete and recreate `destination/` in the project directory. Do not place unrelated files there. If it already contains user data, stop before running the tests and ask how to proceed. +- After a fix, rerun the full suite. For this baseline, the intentionally disabled test may remain skipped. +- If validation is blocked by the environment, report the blocker and distinguish tests actually run from checks inferred from reading code. + +## Working Approach + +1. Read the relevant implementation and tests before editing. +2. Reproduce the reported issue and identify its cause. An existing failing test can serve as the regression test; add tests when meaningful coverage is missing. +3. Make the smallest clear change that addresses the requested task. +4. Validate the result and inspect the final diff for unintended changes. + +For substantial work, briefly state the intended approach. Ask questions when ambiguity materially affects correctness or scope. For routine choices, proceed with reasonable assumptions and state them when they matter. + +## Benchmark Integrity + +- Preserve existing test expectations. Do not weaken assertions, disable tests, or exclude failing tests to obtain a passing build. +- Do not change build or test configuration merely to hide a failure. +- If evidence indicates that a test expectation is wrong, explain the conflict and ask before changing that expectation. +- Do not hard-code behavior for test names, fixtures, or specific test data. +- Keep the intentionally disabled large-data test disabled unless the task requires enabling it. +- Do not add solution hints or the known bug's cause to benchmark instructions or documentation unless explicitly requested. + +## Change Discipline + +- Keep changes focused on the user's request and match the surrounding code style. +- Avoid unrelated cleanup, formatting changes, new features, speculative abstractions, and unnecessary dependencies. +- Remove imports or code made unused by your own changes. Leave unrelated existing code alone. +- Preserve pre-existing user changes. Do not reset, overwrite, or revert unrelated work. +- Do not commit or push unless requested. + +## Final Report + +Briefly explain: + +- What you found and, if applicable, the root cause. +- What changed and why. +- The validation commands actually run and their results, including failures or skipped tests. +- Any remaining limitations or blockers. + +A passing test suite is evidence of correctness, not a substitute for understanding the behavior and reviewing the diff. diff --git a/README.md b/README.md new file mode 100644 index 0000000..923c501 --- /dev/null +++ b/README.md @@ -0,0 +1,85 @@ +# Coding Agent Benchmark + +A small Java project for comparing coding agents. Its compact codebase and existing test suite provide a starting point for observing how an agent understands unfamiliar code, investigates bugs, and makes focused changes. + +This project is intended as a benchmark and learning environment. It does not include a runnable application; start with the source code and unit tests. + +## The Example Project + +`SplitZipWriter` writes files from input streams into ZIP archives, creating multiple parts based on a size threshold. Archives are named `_01.zip`, `_02.zip`, and so on. Each input file is written entirely into one archive; switching archives happens between files. + +Other components: + +- `ZipResult` collects the names and sizes of the generated archives. +- `OutputStreamEnhancer` allows the output stream to be wrapped with additional behavior. +- `InputStreamEnhancer` defines the corresponding interface for input streams. +- `SplitZipWriterTest` checks archive splitting and file contents. + +## Requirements + +- JDK 21, matching the repository's `.java-version` +- A filesystem supporting POSIX file permissions, such as those typically used on Linux or macOS +- Internet access for the first build to download Gradle and dependencies + +The Gradle Wrapper is included, so no separate Gradle installation is required. `JAVA_HOME` should point to the JDK you want to use. + +## Quick Start + +Run all tests from the project directory: + +```sh +./gradlew test +``` + +Run a full build, including tests: + +```sh +./gradlew build +``` + +Force the tests to run again, even if Gradle considers them up to date: + +```sh +./gradlew test --rerun-tasks +``` + +The HTML test report is available at `build/reports/tests/test/index.html` after the test run. + +**Warning:** The tests delete and recreate the `destination/` directory inside the project directory. Do not store your own files there. + +## Using This Project as a Benchmark + +The baseline contains one failing unit test. Another test covering large amounts of data is disabled with `@Disabled`. A failing test run is therefore part of the benchmark's initial state. + +Example prompt for a coding agent: + +> Inspect the project and run the unit tests. Find the cause of the failing test and fix it with the smallest reasonable change. Preserve the existing test expectations and do not disable any tests. Explain the cause, your change, and how you validated it. + +For comparable runs: + +1. Start each agent from the same commit in a separate checkout. +2. Use the same prompt, Java version, and access permissions. +3. Record the model, agent version, settings, runtime, and token usage where available. +4. Review the resulting diff and run the tests again. + +Possible evaluation criteria include a correct root cause analysis, a small and understandable change, preserved test expectations, and successful validation. Passing tests alone do not replace reviewing the solution. + +## Project Structure + +```text +. +├── build.gradle # Dependencies and test configuration +├── settings.gradle # Gradle project settings +├── gradlew / gradlew.bat # Gradle Wrapper +├── AGENTS.md # Guidance for coding agents +└── src + ├── main/java/illgner/ch + │ ├── createzip # ZIP creation and result data + │ └── stream # Stream enhancement interfaces + └── test/java/illgner/ch + └── createzip # Unit tests for SplitZipWriter +``` + +## Technology + +Java, Gradle, JUnit Jupiter 6, Hamcrest, Apache Commons IO, and Apache Commons Lang.