← All posts
Blog

Prioritizing a 238-task to-do list with AI in about a second

My task manager has no priority field, on purpose. So I tried calculating priority fresh instead: an AI model answers a few questions about every open task, and code turns the answers into a ranking. All 238 tasks take about a second and cost under a penny.

My task manager, Things, is wired into my AI setup. Agents groom my list twice a day, file new tasks, and check items off when they finish the work.

What it doesn't have is a priority system, and that's on purpose. I follow GTD (Getting Things Done), which skips "urgent" and "low priority" tags because every day is different. Energy, competing commitments, and whatever lands unexpectedly keep priority fluid.

So I tried calculating priority fresh each time instead of storing it.

Asking a decision model, not a chatbot

The judgment calls come from Jev, a decision model from TypeSafe. Jev doesn't write text. It answers yes-or-no questions and picks from scales, and returns a probability for each answer. For every open task, it answers these:

Question Answer
Is this a real obligation or commitment, rather than optional backlog? yes or no
What happens if it slips two weeks? nothing, minor annoyance, real cost, or serious harm
When does it realistically need to happen? no window, this month, this week, or today
Does it need starting now because of lead time, like shipping or booking ahead? yes or no
Is someone else waiting on it? yes or no
Can I finish it in one sitting, without waiting on anything? yes or no
Does it depend on an earlier open step in its project? yes or no

Each question is a few lines of plain English. Here's the one about slipping:

slip: {
  type: "score",
  instructions: "What happens if `task` slips two weeks from `today`?",
  criteria: [
    "Nothing: no one notices and nothing is lost",
    "Minor annoyance: a small inconvenience, or it just sits on my list longer",
    "Real cost: it costs money, an opportunity or window closes, or I let someone down",
    "Serious harm: a health, legal, financial, or security risk, or a major commitment broken",
  ],
},

What the model sees

A good answer needs context, so each task goes to Jev with more than its title: its notes, its tags and what each tag means, its project and area and their descriptions, and the open steps that come before it in its project. A short note on my current goals goes along too, so the answers can weigh a task against what matters this week.

Dates are spelled out for it. A deadline arrives as "Sunday, October 4, 2026 (in 3 days)" or "OVERDUE", so the model never has to do date math. When a fact can be computed, compute it and hand it over rather than asking the model to work it out.

The ranking is code, not AI

Jev makes the judgment calls. Turning them into a ranking is plain arithmetic:

score = real × (weighted sum of slip, when, lead time, waiting, one sitting, deadline)

"Is this real?" gates everything, so optional backlog sinks no matter what else is true. The rest are weighted, with separate weights for two views:

Signal Today This week
What happens if it slips 0.20 0.35
When it needs to happen 0.30 0.25
Lead time 0.10 0.10
Someone waiting 0.10 0.15
One sitting 0.20 0.05
Deadline 0.10 0.10

The today view rewards what's doable and due. The week view rewards what's consequential. The deadline signal is computed in code, not asked, and a few of my tags nudge the score: a task tagged Waiting drops by half, for example.

Keeping the weights in code is the most useful design choice in the whole thing. The model's answers don't change when I adjust a weight, so I can tune the ranking as often as I like without asking the model anything again.

Fast and cheap enough to rerun anytime

Speed is what makes it practical. Each judgment takes about 150 milliseconds, and I run 32 at once, so the whole list of about 240 tasks is scored in about a second.

It costs under a penny. Jev charges $0.042 per million tokens of input, and a full run comes to about nine-tenths of a cent.

Answers are cached, so a second run takes under a tenth of a second. Today's date is part of what the model sees, so every task gets a fresh look once a day, and again whenever the task or its project changes.

Where it stands

It's still a prototype: a command-line tool that reads my task list and prints three rankings, each task with a breakdown of its signals so I can see why it landed where it did. Next I'm putting it to work in my daily review, to find out whether it actually changes what I pick.

It's the same idea I build for businesses: let AI sort and rank the pile, and keep a person in charge of the decisions. I used the same model to label GitHub issues automatically. The pattern fits anywhere someone faces a long list and has to decide what comes first:

  • A support queue: rank open tickets by how upset the customer is, what's at stake, and how long they've waited, not just by arrival order.
  • Sales leads: score each lead on fit and timing, so a salesperson starts the day with the ones most likely to close.
  • A backlog of requests: rank internal or customer requests by impact and effort, as a starting point for the planning meeting.
  • Collections: order overdue accounts by amount, age, and history, so follow-ups go where they matter most.

In each case, the ranking is a suggestion. A person still decides what to do first.

Questions or thoughts? Discuss this post on LinkedIn.

Have a process that eats your week?

Tell me about it on a free 30-minute call. You'll leave with a clear sense of what's worth automating, whether or not we work together.

Book a discovery call