How do you run performance reviews when some of your team is non-human?
Genuinely working through this in the People playbook. My team is a mix of humans and AI agents running repeating workflows. The human performance review cadence makes sense. But how do you evaluate and improve the agent side? Benchmarking prompt quality? Tracking output variance over time? Curious what frameworks people are using here.