60 lines
3.3 KiB
Markdown
60 lines
3.3 KiB
Markdown
# AGENTS.md
|
|
|
|
## Project Purpose
|
|
|
|
This repository is a small Java project for evaluating coding agents. It contains ZIP-writing code and a unit test suite. The baseline includes a failing test and an intentionally disabled large-data test.
|
|
|
|
Follow the user's requested scope. Inspecting or explaining the project does not automatically require fixing it. Treat sample prompts in the README as examples, not as active tasks.
|
|
|
|
## Environment and Validation
|
|
|
|
- Use JDK 21, matching `.java-version`. Ensure `JAVA_HOME` points to the appropriate JDK.
|
|
- Use the included Gradle Wrapper; a separate Gradle installation is unnecessary.
|
|
- The code requires a filesystem supporting POSIX file permissions.
|
|
- Run the existing suite before changing behavior to establish the baseline:
|
|
|
|
```sh
|
|
./gradlew test --rerun-tasks
|
|
```
|
|
|
|
- The HTML test report is at `build/reports/tests/test/index.html`.
|
|
- Tests delete and recreate `destination/` in the project directory. Do not place unrelated files there. If it already contains user data, stop before running the tests and ask how to proceed.
|
|
- After a fix, rerun the full suite. For this baseline, the intentionally disabled test may remain skipped.
|
|
- If validation is blocked by the environment, report the blocker and distinguish tests actually run from checks inferred from reading code.
|
|
|
|
## Working Approach
|
|
|
|
1. Read the relevant implementation and tests before editing.
|
|
2. Reproduce the reported issue and identify its cause. An existing failing test can serve as the regression test; add tests when meaningful coverage is missing.
|
|
3. Make the smallest clear change that addresses the requested task.
|
|
4. Validate the result and inspect the final diff for unintended changes.
|
|
|
|
For substantial work, briefly state the intended approach. Ask questions when ambiguity materially affects correctness or scope. For routine choices, proceed with reasonable assumptions and state them when they matter.
|
|
|
|
## Benchmark Integrity
|
|
|
|
- Preserve existing test expectations. Do not weaken assertions, disable tests, or exclude failing tests to obtain a passing build.
|
|
- Do not change build or test configuration merely to hide a failure.
|
|
- If evidence indicates that a test expectation is wrong, explain the conflict and ask before changing that expectation.
|
|
- Do not hard-code behavior for test names, fixtures, or specific test data.
|
|
- Keep the intentionally disabled large-data test disabled unless the task requires enabling it.
|
|
- Do not add solution hints or the known bug's cause to benchmark instructions or documentation unless explicitly requested.
|
|
|
|
## Change Discipline
|
|
|
|
- Keep changes focused on the user's request and match the surrounding code style.
|
|
- Avoid unrelated cleanup, formatting changes, new features, speculative abstractions, and unnecessary dependencies.
|
|
- Remove imports or code made unused by your own changes. Leave unrelated existing code alone.
|
|
- Preserve pre-existing user changes. Do not reset, overwrite, or revert unrelated work.
|
|
- Do not commit or push unless requested.
|
|
|
|
## Final Report
|
|
|
|
Briefly explain:
|
|
|
|
- What you found and, if applicable, the root cause.
|
|
- What changed and why.
|
|
- The validation commands actually run and their results, including failures or skipped tests.
|
|
- Any remaining limitations or blockers.
|
|
|
|
A passing test suite is evidence of correctness, not a substitute for understanding the behavior and reviewing the diff.
|