← All posts
Blog

Auto-labeling GitHub issues with AI for under a penny per hundred

I handed one of my least favorite chores to a decision model. It labels every new GitHub issue for less than a penny per hundred, and it flags the ones it isn't sure about instead of guessing.

Between my own apps, I have over 300 GitHub issues to keep organized. Every one is supposed to carry labels:

  • Type: bug, feature, polish, research, docs, or maintenance
  • Priority: how urgent it is
  • Size: how much effort it will take
  • Platform: which device it's specific to, if any
  • Status: whether it's blocked waiting on something

Labels are what make an issue list usable. They let me pull up "every quick bug on iPad" or "everything waiting on Apple" in one click. But labeling is the step I skip when I'm busy, and busy is when most issues get written.

So now GitHub does it for me.

How it works

When I open an issue, a GitHub Actions workflow sends its title and description to Jev, a decision model from TypeSafe, using the Jev GitHub Action. Jev doesn't write text. It answers multiple-choice and yes-or-no questions about whatever you give it, with a probability for each answer.

That's a good fit for labeling, because labeling is really a handful of small decisions:

  • What kind of work is this: something broken, a new feature, polish on something that already works, an open question to research, docs, or maintenance?
  • How urgent is it, compared with other issues?
  • How much effort will it take: an hour or less, a session or two, or several?
  • Does the issue name iPad as a device where the problem happens? (One yes-or-no question per platform, so an issue can carry several.)
  • Does it say the work can't start until something outside it happens?

The questions live in a JSON file in the repository, written in plain English. Here's the one for blocked issues. A noul is Jev's yes-or-no question type:

"status: blocked": {
  "type": "noul",
  "instructions": "Does the GitHub issue in `issue` say the work cannot proceed until something outside the issue happens?",
  "criteria": {
    "true": "It names a blocker that is not resolved yet: waiting on an approval, an entitlement, a third party, another unfinished issue, or a decision someone else has to make.",
    "false": "Nothing stops work from starting now. Background, related work, design references, or steps that come first within the issue itself are not blockers."
  }
}

Writing the criteria is most of the work. Notice that the "false" answer spells out what doesn't count. Related work and "do this step first" are exactly the things a quick reading mistakes for a blocker.

Confidence decides what gets applied

The probabilities are the secret ingredient. A short script turns Jev's answers into labels, and each kind of label has its own bar:

  • Type is always applied, since every issue needs one. If Jev is less than 50% sure, the script also adds a "needs triage" label, so the guess is flagged for me to check.
  • Priority and size go on only at 80% or higher. Most issues are normal priority and medium size, which in my scheme means no label at all, so leaving them off is usually right.
  • Yes-or-no labels like platform and blocked go on at 70% or higher.

It also never touches an issue I've already labeled. If an issue carries any label besides "needs triage", a person has been there, and the workflow leaves it alone. It checks again right before applying anything, in case I labeled the issue in the meantime.

Testing it before turning it on

Before letting it label anything, I ran it in dry-run mode against 97 issues I'd already labeled by hand. A dry run writes the labels it would apply, along with Jev's probabilities, to a report instead of to the issues. That let me tune the thresholds against my own judgment.

It matched my call on the type of work for 82 of the 97 issues, about 85%, and most of the misses could honestly go either way.

Testing also showed me what to leave out, which mattered as much as what to put in:

  • Some questions weren't worth asking. At first I asked about every status label, not just blocked. Jev answered those confidently and wrongly, because on a brand-new issue those statuses are almost never true. High confidence isn't the same as being right, which is why you test against real data. Now the workflow asks only about blocked.
  • Some facts shouldn't be the model's job. Jev kept adding an Android label to issues for an app that doesn't run on Android, just because they mentioned it, because it had no reliable way to know which platforms a repository ships on. Now a repository setting lists them, and any platform outside that list is dropped. That cut the wrong platform labels from 6 to 4. When a fact is already known, put it in the code rather than asking the model to work it out.

What it costs

Each issue is about 2,000 tokens of text, and Jev charges $0.042 per million tokens. That works out to about eight-tenths of a cent to label 100 issues. Building, tuning, and testing the whole thing cost less than a dime.

Where this pattern fits

The whole thing is one workflow file, one file of questions, and a short script. To use it in another repository, I copy those files over and set the list of platforms.

It's also the same pattern I build for clients: anywhere a person reads something and makes a few quick calls. Some ideas:

  • Support email: label each message by topic and urgency and route it to the right person, flagging anything unclear instead of guessing.
  • Sales leads: tag each inquiry by service, size, and fit, so the best ones get a reply first.
  • Customer feedback: sort survey answers and reviews into themes, so the complaints that keep coming up are easy to spot.
  • Invoices: flag the ones that need a second look, like a total that doesn't match the purchase order or a vendor you haven't paid before.

My App Store review monitor works the same way: it sorts each new review and decides whether it needs a reply before anything else happens.

Not every AI job needs a chatbot. A lot of everyday busywork is really a handful of small decisions, and a model built for decisions can make them for almost nothing.

Questions or thoughts? Discuss this post on LinkedIn.

Have a process that eats your week?

Tell me about it on a free 30-minute call. You'll leave with a clear sense of what's worth automating, whether or not we work together.

Book a discovery call