Applied research on agents, automation, and private AI.
We study how AI systems behave inside real businesses — and publish what we learn. What survives the lab becomes a Nexos service.
Programs
Research agenda
Four programs. Publications appear under the program they belong to.
Agent architectures for real workflows
How multi-agent systems should be structured to run actual business processes — delegation, tool use, and hand-offs that hold up outside a demo.
OutputReference architectures + open benchmarks
Approval & autonomy design
When an agent should act, ask, or defer — and how to make human-in-the-loop feel like one clear decision instead of a growing queue.
OutputDesign patterns + a scored eval set
Private & on-prem AI
Making open-weight models genuinely usable in-house: hardware sizing, tuning, and the honest trade-offs against hosted frontier models.
OutputSizing guides + deployment playbooks
Automation reliability
Treating agent automations like production software — tracing, retries, self-evaluation, and the failure modes that only appear at scale.
OutputTooling + a reliability checklist
Artifacts
Open source & artifacts
Everything the lab releases publicly — code, benchmarks, and playbooks.
Workflow-Agent Bench
A suite of real back-office tasks — intake, triage, quoting, reporting — for measuring whether an agent actually completes the job, not just answers.
Read the write-up →Approval-gate patterns
Reference implementations of the human-in-the-loop patterns we deploy: draft-then-approve, autonomy dials, and edit-before-send.
Read the write-up →On-prem LLM sizing guide
How we size hardware and pick an open-weight model for a given workload — the questions we ask before quoting an on-premise install.
Read the write-up →Programs
Ways to work with the lab
Two routes in, without joining full-time.
Research residency
A 3–6 month embedded stint working on one lab program, with a published artifact as the exit criterion. Open to engineers and researchers.
ApplyPilot collaborations
We take a small number of businesses per quarter as live testbeds for lab work — you get early access to what we build, we get real-world evidence.
Pitch a pilotFocus areas
What the lab works on
The lab is organized around disciplines, not headcount — each maps to a program above.
Agent systems
Discipline
Architectures, delegation, and tool use for agents that run real processes — the core of what graduates into the platform.
Private AI & infrastructure
Discipline
On-prem and hosted private models: sizing, tuning, and deployment so a business can own its AI without owning a research team.
Reliability & evaluation
Discipline
Making automations dependable — tracing, self-evaluation, and the benchmarks that tell us whether the work actually held up.
Want to own one? Work with the lab →
About
About the lab
Nexos Research is the applied-research arm of Nexos. Building AI and software inside real businesses puts hundreds of live workflows in front of us a year; the lab turns that vantage point into evidence — testing agent architectures, deployment patterns, and automation designs against real workloads instead of benchmarks alone.
We publish what we learn, release tools where we can, and graduate what works into Nexos services. The lab operates with its own agenda and identity, and will move to its own home (research.nexosai.net) as its output grows.
research@nexosai.net