We've reached a point where LLMs are capable enough to power many agentic workflows, yet relatively few AI agents make it into stable, long-term production.
In your experience, what's been the hardest engineering challenge to solve?
- Tool reliability?
- Long-term memory?
- Planning and reasoning?
- Context management?
- Evaluation and benchmarking?
- Authentication and permissions?
- Multi-agent orchestration?
- Cost and latency?
- Human-in-the-loop approval?
- Something else?
If you've deployed AI agents in production, I'd love to hear what actually broke, what surprised you, and what lessons you learned. Real-world experiences are far more valuable than demo successes.
[link] [comments]