Benchmarking and Advancing Agentic Orchestration
Capstone practicum with Research Net.AI: systematically benchmarking deployed LLM agents that turn research into production-grade software.
Northwestern University × Research Net.AI (industry capstone) · In progress
Research question
How can the performance of deployed multi-agent systems for automated software engineering be measured systematically, and where do their natural language understanding (NLU) pipelines break down?
Methods
- Standardized benchmark suites to quantify agent performance
- Stress-testing of existing agentic architectures
- Analysis of NLU bottlenecks in artifact generation
- Multi-agent orchestration in a production environment