Most AI agents look great in a demo and fall apart in real work. Dave Meyer, VP of Product for Jira at Atlassian, has spent the last three years watching that happen.

At our State of the Agent event, he sat down with Pendo CEO and co-founder Todd Olson to talk about what Atlassian scrapped, what finally worked, and why context decides whether an agent is useful or just expensive. Below is an edited version of that conversation.

If you’d prefer to listen, catch the full conversation here.

Q: What does the future of Jira look like in an AI-native world?

Dave: There are really two halves to it. The first is what we call the autonomous system of record for work. Jira's job has always been the same: it's a system of record for the work happening in an organization, whether that's software developers, marketing people, or HR operations.

The problem that systems like Jira always run into is human psychology. People don't want to write down what they're working on, because they already know it.

That creates a real tension: the system is valuable to the organization and not particularly valuable to the individual doing the data entry. A big part of what we're building is making Jira work item creation faster and more automated, until it's essentially running as ambient infrastructure in the background.

The second half is the orchestration layer for agents. If you think of a business as a set of workflows, increasing portions of those workflows are going to be handled by an agent rather than a person, and that's especially true in software development.

What separates a human from an agent is that humans are context-accumulation machines. Everything we do builds more memory into our own system, even as that memory gets fuzzy over time. An agent doesn't have that. Every time it picks up a task, it starts from scratch.

The next job is twofold:

  • First, the system needs to track human and agent work in the same place, not just human work.
  • Second, as workflows shift toward agents, we need to supply those agents with the context that makes them effective at their piece of the work.

One half is the system of record, made autonomous. The other half is becoming the orchestration engine for the agents doing the work within a team.

Q: You've been building AI into Atlassian's products for a long time. What has that journey looked like?

Dave: The first stage, back in 2023, was the simplest assistive AI: features that went from "wow, this is amazing" to table stakes within months, even though they're still useful.

What turned out to be the bigger bet was our investment in search. We wanted search to cover everything Atlassian knew, so we built data connectors into platforms like SharePoint, Google Drive, Box, and Zendesk early.

If we can make sense of a user story in Jira, we can make sense of a support case in Jira Service Management or in Zendesk. That semantic indexing work became the foundation for what we now call the Teamwork Graph. That's stage two: great search, plus a connective graph that understands relationships between all of those objects.

Stage three started a little less than a year ago, when everyone realized agents needed to be part of the product. We'd already seeded that investment years earlier by letting customers build agents in Studio, but most of what got built then were skills.

The real shift happened around when Claude Opus 4.5 came out in 2025. That's when we accelerated the strategy: you'll have frontier-level agents from Claude and ChatGPT, you'll have more constrained out-of-the-box agents, and people expect to interact with products agentically.

That triggered a full rearchitecture of Rovo, from a question-answerer into something that takes action. Since Jira is the system of record for work, and agents are going to do that work, our job is to stay as agnostic as possible to agents and be the system that makes all of them more effective.

Q: You mentioned context as the make-or-break factor in an agent. How do you approach giving agents a helpful amount of context?

Dave: That's the separation between what the Teamwork Graph is, and how we make it useful.

The Teamwork Graph is a real-time, permission-aware map of the work, goals, code, documents, and activity across an organization. We’ve also invested heavily in enterprise security and permission-aware traversal, so everything an agent retrieves is validated against who's allowed to see it.

Traversing that graph means we can answer questions like who's affected by a change, and which issue is blocking which goal. That leads to a much better answer than a generic model would produce on its own.

Going forward, a lot of the work is extracting the right information and making it as relevant and token-efficient as possible. You can point a frontier model at a task and tell it to go figure things out, but it'll burn a lot of tokens.

We're trying to do the heavy lifting up front: for every work item in the system, traverse the graph, extract the relevant fragments, and attach context from related past work, so when you hand a Jira issue to an agent, it produces a better result with less prompting.

Q: How do you actually measure "better" here? What does good look like at Atlassian?

Dave: Our definition is simple: an agent should produce higher-quality output, in less time, for fewer tokens. Take a typical large organization asking an agent, "What's Project Titan?"

If that company has had three different projects named Project Titan over the years, most agents today will search for every reference they can find and burn a lot of tokens reasoning through mostly unrelated information.

What we're trying to build is the pre-traversal and sense-making that lets an agent know the relevant Project Titan is the one from last year, tied to the person you actually work with, not the one from five years ago. That's the kind of question agents still struggle with, and it's exactly what context is meant to fix.

Q: If a team is building agents and not seeing the value yet, what's the one thing you'd tell them to do differently?

Dave: Think hard about how much data and how many instructions an agent actually needs. It would be convenient to live in a world where one agent can do everything, and I think a lot of the frontier labs would love for you to believe that's already true. Our experience says otherwise, even with Rovo, which can do a lot but can't do everything.

We get better results by building a small number of more constrained agents. Instead of asking Rovo to triage a work item, we built a dedicated triage agent that's tuned specifically to look at an incoming Jira issue and recommend how the fields should be filled out, what status it should have, and who it should be assigned to.

The other side of that same coin is resisting the pull toward "teams of human-and-agent coworkers" with dozens or hundreds of agents running around. I don't love that framing, partly because anthropomorphizing agents feels risky, and partly because I can barely remember all of my human coworkers, let alone fifty different agents, each one for a different task.

That's why I actually really like Pendo's strategy here: one main agent, so people remember it the way they'd remember a colleague, instead of asking "which agent does that again?"

We're not trying to get people to open a chat with the Jira triage agent. We built the automation in, out of the box, doing one thing well, and if a team wants a more powerful version, the platform is flexible enough to swap it in.

If you're building agents, start here

A few things worth taking directly from this conversation:

  • Separate the record from the reasoning. Getting work documented and current is a different problem from making an agent good at acting on it. Solve the first one before expecting the second to work.
  • Measure "better" in concrete terms. Higher-quality output, less time, fewer tokens. If you can't state your definition that plainly, you can't tell if an agent is actually helping.
  • Build several small agents instead of one that does everything. Constrained, predictable agents are easier to wire into existing workflows and easier for people to trust.
  • Don't make people remember which agent does what. The fewer agents a person has to consciously choose between, the more they'll actually use.

Give your AI agent the context it needs to deeply understand users with Pendo for Agents. Get a demo.