3.3 KiB
AGENTS.md
Project Purpose
This repository is a small Java project for evaluating coding agents. It contains ZIP-writing code and a unit test suite. The baseline includes a failing test and an intentionally disabled large-data test.
Follow the user's requested scope. Inspecting or explaining the project does not automatically require fixing it. Treat sample prompts in the README as examples, not as active tasks.
Environment and Validation
-
Use JDK 21, matching
.java-version. EnsureJAVA_HOMEpoints to the appropriate JDK. -
Use the included Gradle Wrapper; a separate Gradle installation is unnecessary.
-
The code requires a filesystem supporting POSIX file permissions.
-
Run the existing suite before changing behavior to establish the baseline:
./gradlew test --rerun-tasks -
The HTML test report is at
build/reports/tests/test/index.html. -
Tests delete and recreate
destination/in the project directory. Do not place unrelated files there. If it already contains user data, stop before running the tests and ask how to proceed. -
After a fix, rerun the full suite. For this baseline, the intentionally disabled test may remain skipped.
-
If validation is blocked by the environment, report the blocker and distinguish tests actually run from checks inferred from reading code.
Working Approach
- Read the relevant implementation and tests before editing.
- Reproduce the reported issue and identify its cause. An existing failing test can serve as the regression test; add tests when meaningful coverage is missing.
- Make the smallest clear change that addresses the requested task.
- Validate the result and inspect the final diff for unintended changes.
For substantial work, briefly state the intended approach. Ask questions when ambiguity materially affects correctness or scope. For routine choices, proceed with reasonable assumptions and state them when they matter.
Benchmark Integrity
- Preserve existing test expectations. Do not weaken assertions, disable tests, or exclude failing tests to obtain a passing build.
- Do not change build or test configuration merely to hide a failure.
- If evidence indicates that a test expectation is wrong, explain the conflict and ask before changing that expectation.
- Do not hard-code behavior for test names, fixtures, or specific test data.
- Keep the intentionally disabled large-data test disabled unless the task requires enabling it.
- Do not add solution hints or the known bug's cause to benchmark instructions or documentation unless explicitly requested.
Change Discipline
- Keep changes focused on the user's request and match the surrounding code style.
- Avoid unrelated cleanup, formatting changes, new features, speculative abstractions, and unnecessary dependencies.
- Remove imports or code made unused by your own changes. Leave unrelated existing code alone.
- Preserve pre-existing user changes. Do not reset, overwrite, or revert unrelated work.
- Do not commit or push unless requested.
Final Report
Briefly explain:
- What you found and, if applicable, the root cause.
- What changed and why.
- The validation commands actually run and their results, including failures or skipped tests.
- Any remaining limitations or blockers.
A passing test suite is evidence of correctness, not a substitute for understanding the behavior and reviewing the diff.