Agent Skills for Gradle Best Practices and Wrapper Upgrades

Agentic Gradle Plugin 1.0.0 is available and includes skills for Gradle best practices and wrapper upgrades.

Table of Contents

Introduction

In August, we told you Gradle was going agentic. Today, we’re shipping an official plugin that contains the first two Gradle skills. Version 1.0.0 is a starting point, and we have a lot more already in progress.

TL;DR:

Quick refresher if you missed the first post. Agents are now Gradle users we have to design for, right alongside the person at the keyboard. A skill is a small package of instructions that teaches an agent to do one job well, so the agent follows a known procedure instead of guessing from whatever it half-remembers.

What did we ship? #

We shipped a plugin containing two skills, located at github.com/gradle/gradle-skills:

Skill What it does Ask it to…
gradle-best-practices Audits your build against the official best practices, writes a prioritized report, and proposes fixes “Check my build” or “Apply best practices to my build”
gradle-wrapper-upgrade Upgrades an existing Gradle wrapper, pins the checksum, verifies the result, and rolls back if the build no longer configures successfully “Upgrade the Gradle wrapper” or “How do I upgrade the wrapper?”

How do I install the skills? #

We recommend installing our plugin, which gets you all of the skills now and every skill we add later.

For example, in Claude Code:

/plugin marketplace add gradle/gradle-skills
/plugin install gradle-skills@gradle-skills

You can also install through skills.sh, which covers Claude Code, Codex, Gemini, Cursor, and any other agent that reads the Agent Skills format:

npx skills add gradle/gradle-skills

The README has everything else, including how to turn the plugin on for your whole team through .claude/settings.json.

Why did we pick these two? #

We asked you. In our forum discussion and Reddit thread, two requests won by a wide margin: “is my build any good?” and “run the right Gradle command.”

The first became gradle-best-practices.

The second started life as a broad gradle-cli skill, and outside of one job, it didn’t help enough to ship. Agents either already knew the rest of the Gradle command line or ignored what the skill told them. The wrapper upgrade was the exception, so we turned gradle-cli into an upgrade skill.

Neither skill does anything an experienced Gradle engineer couldn’t do by hand. We know. The skills give an agent the right solution instead of a plausible one from a five-year-old Stack Overflow answer.

What does gradle-best-practices do? #

You point your agent at a build and ask it to check it. The agent reads your settings files, build scripts, gradle.properties, wrapper config, version catalog, and anything under buildSrc/ or build-logic/. Then it works through the best practices bundled with the skill and writes a report, with every finding linked to the exact section of our documentation it came from.

You get two modes:

  • Audit (the default): write the report, stop, and offer to fix things.
  • Apply: propose and make the changes, highest priority first.

The first version of our best practices skill embeds the best practices as references instead of fetching them from our documentation. We found that some agent harnesses use a small model to summarize documentation pages before the main model sees them, so two runs against the same docs could come back with different summaries. Including the content of the best practices in the skill itself improved results for some models and made them more consistent overall.

How much does gradle-best-practices help? #

Each model got four small Gradle projects seeded with violations, plus one clean project straight out of gradle init to catch the skill inventing problems.

We ran evaluations across seven models with and without our skills in September 2026. These evaluations covered 40 of the best practices:

Model Without skill With skill Change
Claude Haiku 4.5 11 / 51 29 / 51 +18
Claude Opus 5.5 45 / 51 48 / 51 +3
Claude Sonnet 5 30 / 51 50 / 51 +20
DeepSeek V4 Pro 29 / 51 51 / 51 +22
Gemini 3.8 Flash 38 / 51 46 / 51 +8
GPT 5.6 Luna 27 / 51 47 / 51 +20
Qwen 3.6 35B (local) 11 / 51 21 / 51 +10

Every model improved, and none of the models were perfect without our skill.

Checks passed with and without gradle-best-practices

We saw the biggest gains with DeepSeek-v4-pro, which finished with a perfect score with our skill. Frontier models like Opus 5.5 had less room to improve, but our skill made cheaper models perform just as well. Sonnet 5 with the skill outperformed Opus without the skill, and Haiku 4.5 with the skill achieved a similar score to Sonnet 5 without the skill.

What does gradle-wrapper-upgrade do? #

Upgrading the Gradle wrapper looks like a one-liner. That’s exactly why agents get it wrong.

The skill walks your agent through six steps:

  1. Read the current version and distribution type from gradle-wrapper.properties.
  2. Look up the upgrade version and its SHA-256 checksum on services.gradle.org.
  3. Run ./gradlew tasks to prove the build configures successfully before anything changes.
  4. Run the wrapper task twice.
  5. Check that the appropriate wrapper files changed, then run ./gradlew tasks again.
  6. If the build no longer configures, restore the wrapper, confirm the restore, and tell you what broke and which version to try next.

Why run wrapper twice? Your first wrapper run still uses your old Gradle, so it points the properties file at the new version but regenerates gradlew, gradlew.bat, and the wrapper jar from the old version’s templates. The second run uses the new Gradle and finishes the job. Skip it and your properties file points to 9.0.0, but your scripts and wrapper jar are still from 8.5.

The skill also teaches your agent the difference between “upgrade the Gradle wrapper” and “how do I upgrade the Gradle wrapper?” The first means “do it.” The second means “show me the commands to run”.

How much does gradle-wrapper-upgrade help? #

We ran several evaluations per model with and without our skill, including two where the upgrade can’t succeed:

Model Without skill With skill Change
Claude Haiku 4.5 10 / 29 26 / 29 +16
Claude Opus 5.5 16 / 29 29 / 29 +13
Claude Sonnet 5 12 / 29 28 / 29 +16
DeepSeek V4 PRO 13 / 29 29 / 29 +16
Gemini 3.8 Flash 15 / 29 29 / 29 +14
GPT 5.6 Luna 13 / 29 29 / 29 +16
Qwen 3.6 35B (local) 12 / 29 17 / 29 +5

Four models (Opus 5.5, GPT-5.6 Luna, DeepSeek V4 Pro, Gemini 3.8 Flash) achieved a perfect score with our skill. Our skill improved cheaper models like Sonnet and Haiku to near-perfect scores.

Without the skill, models were able to partially perform an upgrade, but they all failed to pin the Gradle distribution checksum.

When an upgrade failed, none of the models rolled back to a known good state without the skill.

We found that our skill reduced the number of tokens used, which we didn’t expect. Agents without the skill tried many things: they read the wrapper jar, grepped for version strings, hand-edited the properties file, and retried repeatedly. Agents with our skill just followed the upgrade procedure and stopped.

How did we test our skills? #

Our evaluations are driven by an internal harness built on top of Inspect AI. We give each evaluation one Gradle project, one prompt, and two runs that differ in exactly one thing (skill or no skill).

We run our evaluations this way:

  • Every run is one-shot. The agent gets a prompt and nobody steers it after that.
  • Every check is deterministic. Scripts do the grading: they look at the final project, run behavioral probes against the build, and pattern-match what the agent said. We do not currently use an LLM-as-a-judge workflow.
  • Every run happened once (n=1).
  • We measured the noise by running the baseline multiple times.
  • The test projects are private for now. The scenarios, fixtures, and scorers live in a separate private repository (this helps prevent models from cheating by looking up the answers). Our evaluation reports are public.

We published evaluation reports for each skill in the repo’s evals/ directory. This format is more useful to us than to you, and it is very much AI-generated. We are developing a more readable, more human format to use in the future.

What’s next? #

We’re not done with these two skills. We have a growing list of improvements to try, and our evaluations let us test each one instead of guessing.

  • 1.0.1, soon. Our own pre-release review caught the two skills giving contradictory advice about upgrading Gradle, plus wording that could let an unattended audit apply structural fixes. We’re fixing both, re-running the evals, and publishing what changed.
  • More skills. A configuration cache compatibility fix skill is already being tested.
  • Better docs for agents. We want Gradle’s documentation to be something an agent can reliably retrieve from our docs site. Once it is, gradle-best-practices can point at the live docs again instead of carrying its own copy of the catalog, so the skill never falls behind a docs change.
  • A closer look at our evaluation tool gradle-eval, in its own post.

If you run these skills, we want to hear what worked and what didn’t. Found a case we didn’t test? Even better. Come tell us in #agentic-gradle on the Gradle Community Slack. Your reports decide what goes into the next release.

Discuss