Skip to content
AIMUG
← All talks

Mon, Oct 5, 202614:25EvalsMCPCoding agents

Harbor evals for coding agents

Share this talkLinkedInX

Jeff Linwood

Jeff Linwood shows how to test if agents use MCP tools with Harbor Framework.

Watch on YouTubeRecorded at AIMUG, Oct 5, 2026Recap of the nightSlides and notes

Takeaways

  • Set up repeatable tests for agent behavior using the Harbor Framework.
  • Verify if an agent used specific MCP tools by inspecting the audit log.
  • Determine that short instructions in agents.md perform better than long descriptions.
  • Run multiple trials to account for the non-deterministic nature of LLMs.

Resources

Transcript

144 lines

Captions generated automatically from the recording. They may contain errors. Suggest a fix

0:00Intro to Agent Evals

  1. 0:04you All right, so I am going to talk to all of you about… Yeah, let me grab this.
  2. 0:12I am going to talk to all of you about setting up evals for your coding agents.
  3. 0:17So this is going to be a very useful thing for all of you to know.
  4. 0:21If you've ever wanted to instrument your Claude codes, your codexes, possibly, you know, your open codes, anything else out there that's really much more of an agent harness.
  5. 0:33So think about this, not so much for individual models, but if you want to figure out, should we be using Claude code with Opus, who are using Asana, etc, and you actually wanted to get some real numbers on that,
  6. 0:46I'm going to show you how using something called HARP.
  7. 0:50So, in my case, I'm building this application here, RobotBridge, and what it is, is it's basically in this form.

0:54The MCP Tool Problem

  1. 0:58It's a task board that uses MCP to talk to various coding agents.
  2. 1:03However, the problem I ran into was that the coding agents were not picking up the tools that it was providing over MCP.
  3. 1:12So I had to go out and solve that problem.
  4. 1:17What was kind of interesting about this was that, you know, I could tell the MCP connection was there.
  5. 1:24The tools were available, but when I asked that coding engine to pull a task off the board, it just skipped all that, moved straight to prep, and tried to figure out what to do, right? So it wasn't picking up on my MCP tools.
  6. 1:39And I wanted to solve that.
  7. 1:41Now, the FIPS is going to be a couple of lines in your agents.md file, or a couple of lines inside of your claw.md file, but… what's the right wording for that, right?
  8. 1:53Because we have a lot of different things we could put into an agents.md or into a Claude.md.
  9. 1:58Has anyone here tried to instrument that, or ever explore how much, or what the right thing is?
  10. 2:04Okay, well, we've got one person, excellent, so perfect.
  11. 2:07I'm gonna show you a way that you could actually go through and figure that out on your own, and run a few experiments, yourself.
  12. 2:15Because that's what I do.
  13. 2:16I use something called an eval, which we've heard about a few times today already.
  14. 2:20For those of you that are a little bit newer, maybe to LLMs and everything else, what we're going to do is we're going to set up some repeatable tests for behavior.

2:29Harbor Framework Setup

  1. 2:29And in this particular case, I'm working with a tool called Harbor Framework, which is an open source tool.
  2. 2:35It's a Python test harness, so it doesn't matter whether it's Python, it could be anything.
  3. 2:39And their terminology is going to be that we have a task.
  4. 2:43So, we have a realistic prompt and a good starting state.
  5. 2:46We have an agent that is under test. In my case, my tool is a macOS agent workspace.
  6. 2:51Right now, it works with Claude Code, and it works with Codex, and so those are the two that I was exploring.
  7. 2:58You also pair it with a model, because you're going to get very different performance from something like Sonnet than you would with Haiku or with Opus.
  8. 3:07And again, you can explore all that as part of your experimental design.
  9. 3:12Next, we need to have some kind of verifier.
  10. 3:15So, just shipping off a task over to an agent, how does that actually know that it got run, right?
  11. 3:21So, how do we verify that the task actually happened?
  12. 3:24And so, in my particular case, I was measuring two things.
  13. 3:28Did the coding challenge actually work?
  14. 3:31In other words, did the coding agent implement some random task?
  15. 3:35And then second, did it use the tools off my MCP to actually do the task, right?
  16. 3:41And that's two separate things you need to verify.
  17. 3:45And then the last thing that you really have to consider is trials.
  18. 3:50We all know these large language models, they're non-deterministic, so you can't run it once, you have to run it multiple times. And unfortunately.
  19. 4:01every time you run it, it's gonna cost you more money, or use up more usage, and so, you know, you do need to figure out, are you working with a giant budget, or are you working with, like me, you know, a very small project because,
  20. 4:12you know, or a very small budget, because I'm trying to figure out a very small problem.
  21. 4:19So, the tool I used is called Harbor. Does anyone use that one?
  22. 4:24Okay.
  23. 4:25And the reason I bring this up is because I'm using it to evaluate whether or not my tool calls come out through MCP, whether they get picked up, but you could use it for all sorts of evaluations.
  24. 4:37So, there's been some people talking about, I don't know if we are using the right models for our coding tasks.
  25. 4:43Or perhaps you're doing something where, like Colin mentioned.
  26. 4:47or you know, you're pulling the traces down of your large language models.
  27. 4:52Now you actually have some real-world tasks that you have worked on.
  28. 4:56You could take some of those and actually benchmark those against various agent harnesses, against various models, and figure out, hey, maybe I can get by with, like, Luna, right, or Terra. I don't have to use the big ones.
  29. 5:10So, just know that Harper Framework's out there. It's all open source.
  30. 5:14It was really easy to set up.
  31. 5:16I'll be honest, it was really easy for Claude Code to set up for me.
  32. 5:19But, the other thing is that I used it locally with Docker, just running on this regular old Mac.
  33. 5:25You can also host it on a cloud-provided sandbox provider, like Daytona.
  34. 5:29If you want to scale up and above and beyond what your desktop or laptop can do.
  35. 5:34Another thing I'll mention here, I happen to use API… I use an API key with both Claude Code and with Codex, because I didn't want to run into usage limits, and then I'm actually not really sure what what the terms of service are,
  36. 5:48if you're using this sort of thing with your subscriptions.
  37. 5:52So, I just used API keys.
  38. 5:54I did not spend a lot of money on this experiment, but it also limited my scope of the experiment, so I didn't spend a lot of money.
  39. 6:02So, just know that there is a way to set both of these up with subscriptions.
  40. 6:07I just set up a .env file and let it go to town.

6:13Task Structure and Verifiers

  1. 6:13When you set up Harbor, each individual task that you are going to benchmark is a folder, and so you're going to have an instruction.md That's gonna be the actual prompt the agent gets, and then you're gonna have an environment with the repo.
  2. 6:28In my case, what we're really testing is something I've built called Zebric, which is an agent-native application framework.
  3. 6:35That's what's providing the MCP server and the tools associated with the tasks, so it pulls that all together, and then it's also got the test to actually verify that the thing ran. So… My, my framework has an audit log,
  4. 6:49so it can actually tell what agent called what tool, and so it can inspect the audit log and just see if the tool actually got called.
  5. 6:57That makes it relatively easy with this particular task.
  6. 7:00Otherwise, mostly people are using this to do things like coding tests, so they're, you know, can this application actually… or can the Ernest actually add some functionality to an application, or fix a bug?
  7. 7:12And so you could have a Python script that, for instance, tested the bug.
  8. 7:19There is some prior research about this.
  9. 7:22There's nothing in particular that's exactly like what I'm doing, but I did find 3 different PDFs on Archive that you could actually go in and explore if you wanted to learn a little bit more.

7:34Experiment Design

  1. 7:34My experiment was really, do agents use an available workflow tool when nobody tells them to?
  2. 7:44And we had two questions we were trying to figure out.
  3. 7:48Did it write the code? In other words, did the problem actually get solved?
  4. 7:51And then number two, did it use the MCP tools?
  5. 7:55So the idea here is we've got this task tracker, we're running inside the, bridge here, and we're just trying to figure out everything's connected.
  6. 8:05In my case, I ended up with 120 runs for my first wave of the experiment, so I should say that this was just the very first exploration.
  7. 8:14I did go off and explore a couple different variants down the road, but this is sort of the initial experiment.
  8. 8:21I start with 5 different fixtures.
  9. 8:23Two agents, so I had Claude Code and Codex, because both of those were things that I use.
  10. 8:28Two model tiers, so I did only go with small and medium.
  11. 8:31That was mostly because I wanted to cut costs a little bit.
  12. 8:35And if I wanted to expense a little bit, I could certainly have it use Opus or Solve.
  13. 8:40And then I tried three different conditions, because I thought that's really the most interesting thing.
  14. 8:47My conditions were the control.
  15. 8:49The controls have nothing in your agents.md telling it about the tasks.
  16. 8:54A short instruction.
  17. 8:56So based on some of the research, a short instruction is actually much better.
  18. 9:01Put a very short instruction into Agents.md.
  19. 9:04And then the third was the product construction that's currently inside the Mac app, that is live. And so that's the one I was seeing.
  20. 9:11It wasn't really performing as well as I thought it would be.
  21. 9:15But again, I'm just seeing that from user behavior, from actually exploring the app.
  22. 9:20How do I actually boil this down into results?
  23. 9:24Well, I do need to have some sort of ground truth, and so I can look at whether the MCP was used or the task was discovered.
  24. 9:32By looking at the audit log, I can look at whether a task was pulled off the board, because that ends up in the, final ZBERC DB, so it's just a… it's a SQLite database in this case.
  25. 9:44And then I can also check the coding success of the PyTest.
  26. 9:47That's actually something Harbor does for you.
  27. 9:50And then last, if there's any kind of errors, things crash out, that would make an invalid run.
  28. 9:56I did not use Jev, unfortunately.
  29. 9:58I started this project probably about when Jeff came out, so… Okay, here's the results from that 120, run.

10:05Results and Analysis

  1. 10:06The interesting thing, I thought, was that all 120 runs passed the coding tests.
  2. 10:11These were not huge tests, but it was cool to see that, you know.
  3. 10:16That all these things worked.
  4. 10:18What didn't work was the control with nothing in agents.md or Claude MD to tell it to use the tasks.
  5. 10:26The task tracker? It didn't use it.
  6. 10:28It tried to use grep or something, or just ignored it completely.
  7. 10:31With the short instructions.
  8. 10:34I actually got great results from the medium, models.
  9. 10:38So I got 10 out of 10 from… Claude Code Medium, and 8 out of 10 for Codex Medium.
  10. 10:43But the product instructions, what I actually had baked into the Mac app right now that you can download, not so hot, right? I've got, like, 3 out of 10, and 2 out of 10.
  11. 10:54Yeah, or even 0 out of 10 on Codex Media.
  12. 10:57So… This is what those instructions actually looked like.
  13. 11:03On the left-hand side, you can't read it too well, but you can see how short the instructions are. That's the one that worked.
  14. 11:11On the right-hand side is the Mac app instructions, and that's just, like, part of it. It went through and gave a big description about each tool.
  15. 11:20Turns out none of that's needed for effective, tool usage.
  16. 11:24So, you know, with the next version bump of the Mac app, it's gonna get the short instructions.

11:33Key Takeaways

  1. 11:33Some of the key takeaways that I had from this experiment, well, number one was just kind of diving into Harper in the first place and figuring out that, hey, we can start instrumenting our agent harnesses with this,
  2. 11:44evals framework and exploring a little bit more about results.
  3. 11:48Number two, small, medium, small models were worse than medium.
  4. 11:52Again, we're talking about, like, Haiku and, and… Luna, they're not the… they're fast, cheap, not as good, understandable.
  5. 12:00And then shorter instructions performed well, so that was also nice.

12:07Applying This to Your Work

  1. 12:07Now, for all of you, right, because I had my particular problem that I went out to go and solve and actually run some evals on, what could you all do, right?
  2. 12:16So, you know, where could you take this?
  3. 12:19Number one, I think, we all have agents.md files, we all have a lot of .md files in our repos.
  4. 12:25go in and see if those changes in those files actually change how well your agent performs, right?
  5. 12:30That's probably the first thing I'd say if people were interested in exploring this.
  6. 12:35Number two, if you're working on some sort of project that has MCP tools in it, you can actually work on design around that.
  7. 12:42So, in my case, right, like, how are the tools named, what descriptions do you provide? All very useful.
  8. 12:49And then, you can also use it to create your own benchmark if you want to have a state of where your company is at in Agentic Coding now.
  9. 12:59Okay, I'll just leave some links up here.
  10. 13:01The first two sets of links are to my open source projects, which you're welcome to explore, and then the third one is Harbor, which is the eval streamer, which is also open source.
  11. 13:11It's by the people behind Terminal-Bench.
  12. 13:15Any questions?
  13. 13:19Yeah, just… Yes.
  14. 13:25Did I use Prompt Guru or something?
  15. 13:30No, I did not. Haven't heard of it, actually.
  16. 13:34Got it.
  17. 13:35Bye-bye.
  18. 13:36And… Okay.
  19. 13:38Yeah, it was, fascinating and, you know, counterintuitive to what I initially designed in the Mac app.
  20. 13:44So yeah, that's why I thought it was a very interesting client.
  21. 13:49Yes. Not the vertical, but… Colors?
  22. 13:57My color scheme doesn't show up? Oh, no, okay.
  23. 14:02We'll… we'll make sure to put it on the Discord.
  24. 14:05Yeah. You can just send it and link to your presentation. Yeah, sure.

Same night

Transcript20:58The coding factory: canonical specs and sequential agentsJake Cukjati · Coding agents
Premieres Mon, Oct 1216:20Grokbot and Jev, a decision modelJoseph Fluckiger · Agents
Premieres Tue, Oct 1315:01Corvic AI: graph analytics and dependency tracingJames Coffey · Graphs
Premieres Wed, Oct 1425:21Mixture of models and semantic routingColin McNamara · Routing

Related talks