Mon, Oct 5, 202620:58Coding agentsAgents
The coding factory: canonical specs and sequential agents
Jake Cukjati shows how specs, not prompts, let agents code for 23 hours straight.
Watch on YouTubeRecorded at AIMUG, Oct 5, 2026Recap of the nightSlides and notes
Takeaways
- You can get agents to work for hours, with a maximum run time of 23.8 hours.
- Prompting into code is not good, but specs to actual work is the effective method.
- Agents work sequentially one after the other with handoffs to communicate with each other.
- The key part is like one agent at a time during implementation.
- After review, we generate reference materials that feeds into the application so agents learn.
Resources

Transcript
212 linesCaptions generated automatically from the recording. They may contain errors. Suggest a fix
0:00Introduction to AI Agents
- 0:04Hey guys, so my name is Jake Cukjati.
- 0:07I've been pretty much interested in AI agents and Claude Code for the better part of a year.
- 0:15It's been a great discovery, and yeah, I mean, I'm just amazed by what you can do with it.
- 0:22So today I'm going to be talking about my, like, journey from Just taking agents and coding to creating workflows, and how… I, you know, get a coding factory going.
- 0:38So, if you're initially, like, starting out, maybe you're not getting alignment with your coding agents, and you're just telling them, go and create some code, you know, go do XYZ, whatever it may be.
- 0:54But, if you're not getting your agents to run for hours, and you're just sitting idle for, like, 30 minutes at a time.
- 1:03maybe you want to consider how you're actually, you know, performing work with your agents.
- 1:08How can you get them to go, you know, much longer, right?
- 1:12You don't want to be a bottleneck, and that's what I was trying to attempt to do.
1:18Specs Over Prompts
- 1:18So, are you, like, still prompting your agents to code?
- 1:22Because I just put them in a slot, and you just go after it.
- 1:26I can get agents to work for hours, and the most I've gotten is about one day, 20… 23.8 hours.
- 1:37Not quite.
- 1:38Okay.
- 1:39And you can see like here I have like over 112 agents working for 16 hours.
- 1:46Producing 60,000 lines of code. The previous one was 110.
- 1:53So, coding for prompting to code is not good, but specs to actual works.
- 2:00work.
- 2:01That's what you want to do.
- 2:03That's how I found a good medium.
- 2:07getting agents to do everything that I want to do.
- 2:11So I will spend a long amount of time making out specs.
- 2:16And then from the specs, I can just be out to work 40 specs, 20 specs, whatever amount of specs I'm gonna give the agents, and they'll just work all night long.
- 2:28And the good thing about the specs is it's canonical, so you have something to actually verify.
- 2:36I can't stress that enough. Like every single time that there's a bug, the agents figure out how it should work, just because it's like, oh, there's… there's the specs.
- 2:47Like, that's… that's what I'm gonna use to make it actually work.
2:52Dynamic Workflows
- 2:52So maybe you remember a couple months ago I said hey I have this open source project for CTL.
- 3:01That's pretty much obsolete now, because Claude Code came out with dynamic workflows a couple weeks later.
- 3:10And that's what I've been using ever since it's come out.
- 3:16You can replace that, like Agent SDK, or if you use Pi, or any other harness to have those notes.
- 3:23Same thing applies here.
- 3:25We're just calling the, TLIs from, like, shelters.
- 3:30So… It's not just… Like a school line, but… We're not just producing like one thing at a time.
- 3:41Or, you know, multiple cars in that case, when you think about code.
- 3:46this.
- 3:47It's, you know, it's a Kodi factory, and it… Works online badges.
- 3:53Or, you know, like, you can only have one worker at a time. I think… the mythical.
- 3:59Mammoth as a huge role.
- 4:03You can't, I don't use multiple agents working at once to get a job done.
- 4:09They work sequentially one after the other with handoffs. And that's how they do it.
- 4:15communicate with each other.
4:16The Coding Factory
- 4:16you so… There are multiple steps here.
- 4:21We have to start out with a great number of specs.
- 4:27gonna implement.
- 4:28And then we go.
- 4:29champion.
- 4:29meditation.
- 4:31After implementation, there's integration testing.
- 4:34Code red.
- 4:34use.
- 4:35huge And then, triage pretty much to, continue.
- 4:41With, Yeah, what have we learned from this text? What's loose? What wasn't fixed?
- 4:49What's, still open for our continued work?
- 4:53And then, you have to compound on top of that.
- 4:56So, I do want to stress that you spend a lot of time on code review, and with, evaluating my agents.
- 5:05Where I will study their prompts, if they are performing at the efficiency that I want them to, if they're sending too many tokens, too much context window size.
- 5:18I wanted to prevent all that.
- 5:21So the first, like, plugin that I have is Spectacular, where we can dish out our specs.
5:25Tools and Plugins
- 5:28There are multiple types of specs that, that it had taken control over.
- 5:34And basically, it claims how the application will be run, right?
- 5:41The great thing about this fence, like I said before, is it's all verifiable.
- 5:47And then I put this to Forge, which is the workflow, where the agents take it from the spec.
- 5:55To, planning the implementation.
- 5:58Where there's, implementation references, and that's a huge thing that gets built during the compounding, just before we start off with the next, spectrum.
- 6:10But yeah, the agents go through and test everything, implement it.
- 6:16Testing is only, like, concerning.
- 6:19With, unit testing here.
- 6:23The key part is like one agent at a time.
- 6:28And then, after we implement, I use Rigor, which is a skill plugin for.
- 6:35Test harnesses.
- 6:38So, the agents will verify that the specs are done, they'll have a checklist of what they want to implement for integration testing, and, integrate… yeah, do end-to-end testing that way.
- 6:54This works real well for back-end services.
- 6:58I'm not really good with UI, but yeah, back end, it's… It'll get you where you want to go.
- 7:06And like I said before, they have the specs.
- 7:09So if nothing fits the spec, if there's a bug, they'll catch it.
- 7:15So that's awesome.
- 7:16So I'm always in complete control because I just have to evaluate the specs themselves.
- 7:23This isn't running your agent on an infinite browser or anything like that.
- 7:30So yeah, I haven't spent a lot of time with code review.
- 7:36I want to make sure that the quality of the code, the intent, is good.
- 7:41You know, code architecture, infrastructure architecture.
- 7:45That's aligned with what I want to achieve.
- 7:48And then, I always like to ask questions.
- 7:51And I think… Findings in 18 of 18, that's just some random number that my agent found from a previous test run.
- 8:01So, I'm glad I pulled up there.
- 8:06Yeah, so compounding.
8:07Compounding and Triage
- 8:07So after review, we generate reference materials that goes and feeds into.
- 8:14the.
- 8:14application.
- 8:16That's pretty much how the agents learn.
- 8:19That's how I learn how the agents learn, and that's how they figure out.
- 8:27How the system works with these reference materials.
- 8:32What's the patterns organized twice?
- 8:36We want to solidify it.
- 8:38We want to keep on using those good patterns, and then the agents will maintain all that.
- 8:47So triage, just every single time I run a spec, my agents will create two items.
- 8:54If there's any, remaining.
- 8:56Like, to do work that we didn't cover, that's not covered in the specs themselves.
- 9:02So anybody a Warhammer 40k fan?
- 9:08So I run everything based off Warhammer 40K.
- 9:13And I actually use Jeff, which Colin, you gave me that inspiration. That's my Jeff.
- 9:21Product manager… It goes out and, Puts priority on all the tickets, and puts lockers as well.
- 9:35So yeah, this is, pretty much my tarot harness, my Agentic wheel.
- 9:40Creates the entire circle. So, every single run, it gets a little bit better.
- 9:46It's, yeah, it's pretty… It's hard to get started, but after a large amount of work.
- 9:54You can… You can start being real successful.
- 10:04Yeah, and the good thing, like, every… Every run, the prompts get way better, right?
10:14CLI and Context
- 10:14And another thing, maybe this will help Jeff, you were mentioning.
- 10:19you.
- 10:20Evals?
- 10:21Well, one thing I've noticed is if I create a CLI for my agents, and the CLI is the prompt engine.
- 10:29And it has a CAD machine in the COR. It's really, it's really helpful.
- 10:36That's how I can get the right context to the agent at the right time.
- 10:41and.
- 10:41The COI just says, "Hey, come back."
- 10:44You need to call me again.
- 10:46I think… This is another skill that I have that's really helpful.
- 10:55That's… that's pretty much it. But yeah, I just, I want to show this, and… Yeah, keep your sex as a canonical truth.
- 11:07Major sales into CLIs, and then compound everything.
- 11:14I'm gonna open source my, skills.
- 11:17Soon.
- 11:19Not yet, but once I.
- 11:23Get rigor working the way I want it to. Yeah.
11:30Q&A and Scaling
- 11:30What questions do we have for Jay?
- 11:33You're only doing one feature at a time instead of a bunch of spreaders. Yep.
- 11:39One, one inch at a time.
- 11:41using handouts sequentially.
- 11:44Or, JSON.
- 11:47Whatever the output.
- 11:48Started or start pressing back.
- 11:52Yeah, when I started with, 4CTL, because… And Profit said they weren't gonna allow you to use Agent SDK, for the subscription plan.
- 12:04I don't know if you remember that, but they did warn that they were going to prevent that from happening.
- 12:09Like, I knew at that time I could automate it.
- 12:13And the, you know, either calling the COI.
- 12:19Or calling Agent SDK was, like, the best way to do that.
- 12:24But they added dynamic workflows, and I was like, this is great.
- 12:28So, yeah, the first iteration for CTO, just had one agent, and the agent would run to, like, a million, you know, it would hit its contact, context window, compact, and keep on running.
- 12:44It, you know, so there wasn't a lot of efficiency there.
- 12:48So, like, if I show this.
- 12:53Here I have like all the agents. You can see like the The most an agent ever goes on this context window is about 330.
- 13:03That's, like, the highest context window, but I'm trying to optimize this as much as possible, so I have another agent that, you know, evaluates this.
- 13:13And we try to, converge on, like, 250… As our, our soft limit for contracts with those.
- 13:26Any other questions?
- 13:28Right.
- 13:29Julian on thrift.
- 13:32Make an entire set of database tables for me.
- 13:36It's changed the part so much. What would get rid of those same things?
- 13:40What would a great fabric?
- 13:42What would keep it from getting?
- 13:45Bill was slow.
- 13:48So yeah, nothing gets implemented until you run the coding engine.
- 13:54To run the coding engine, you need specs.
- 13:57And to create specs, You evaluate it with the agent.
- 14:02So, if the agent's like, hey, this is a gap.
- 14:06we need to fix this, then you're like, yeah, fix that, get rid of those cables, we don't need them anymore, throw them away. It's old stuff.
- 14:15And the best part is, like, while you're planning it, you didn't code a thing.
- 14:20Not one thing was coded.
- 14:23I mean, if you're… You know, doing, like, long number of sessions, like.
- 14:29You know, like, 40… Specs at once. Like, I can run for, like, 10, 15 sessions.
- 14:39Which is about, you know, like, 20… 20 hours of work.
- 14:45Zoom.
- 14:46change a spec, or would you create a new spec that seeds it, and then you seed it built in between?
- 14:52So… Yeah, as the… okay, so… There's a couple things on that.
- 14:58One is… The… Monorepos.
- 15:06And how your codebase changes over time, and how you handle domains, right?
- 15:12So, domain-driven design is very important.
- 15:16I don't have that problem right now, but I knew I'm gonna… hit it very soon, because my… the Copic side run is approaching 200,000 lines.
- 15:29And… I need more domains.
- 15:33So I have, like, 80 or 90 text files.
- 15:37And each one… So, each spec contains a topic of concern, which sets it apart from all the other specifications.
- 15:48And if you think about it like entry points from like, these are the entry points of this application.
- 15:56Each entry point is spent.
- 15:59can… Like, top down, it's basically like, you know, code, how it maps, and there's references, and.
- 16:07All that, right.
- 16:10So, at a certain time, the amount of context that an agent receives could be exhausting to itself.
- 16:19So… Building a new domain, you can separate your code.
- 16:25you can separate your specs and isolate it to that one new domain, right?
- 16:32So, like, for instance, I'm gonna have to split my, codebase to 3 or 4 domains, right? So if I do that equally, that's about 20 specs.
- 16:4320 specs, the agents can easily handle.
- 16:45Each spec is less than 1,000.
- 16:49Normally, it's around, like, 400 or 500 lines.
- 16:52So you have to be heavily involved.
- 16:55Texture. Oh, yeah.
- 16:57Yes.
- 16:58Yeah, everything gets architectured. I have… Skills for… design HTML files to show architecture, design, because we're going at such a fast velocity.
- 17:14Another thing is, like.
- 17:17computation efficiency.
- 17:19So I do a lot of, like, data pipelines and derived data, like, you know, for data analytics.
- 17:29So, Ronnie skewed as much as I can with… The limit of RAM that I have is what I try to maximize.
- 17:41Yeah.
- 17:44Other questions?
- 17:47Yeah, I'll ask a final question.
- 17:49So you mentioned with your specs that you make specs canonical.
- 17:54So as your different agents are improving the architecture, and they I assume updated a root spec.
- 18:03What's the mechanism that that triggers that canonical reference down to the existing code?
- 18:09Sure.
- 18:46Let's see…
- 18:53Maybe… let's do this one.
- 19:09Okay, and… Let's do…
- 19:20I think it's this one. Yeah, okay.
- 19:26So to keep my specs in line, I use a file. It's called a manifest.
- 19:35What that file is, is it gets generated based off of this one commit.
- 19:41Which contains all these specification files.
- 19:45These are… it's just the depth of all the specs themselves.
- 19:51It uses this one file to outline all those specs.
- 19:56Right.
- 19:57At the very top, it says what the commit hash is, and the agents will use this commit hash with this command, and then… to identify, like.
- 20:12All the specifications for that batch.
- 20:15This gets split up into multiple batches by a scheduler.
- 20:20Which.
- 20:24Should be able to show right here, implemented badges.
- 20:30And each one of these contains the hash of the canonical commit, the, the files that it uses, what it's blocked by, such and so forth.
No lines match.
Same night
Transcript14:25Harbor evals for coding agentsJeff Linwood · Evals
Premieres Mon, Oct 1216:20Grokbot and Jev, a decision modelJoseph Fluckiger · Agents
Premieres Tue, Oct 1315:01Corvic AI: graph analytics and dependency tracingJames Coffey · Graphs
Premieres Wed, Oct 1425:21Mixture of models and semantic routingColin McNamara · Routing

