The system uses intelligence to manage intelligence. Routing models estimate which capability can satisfy a request under quality, latency, privacy and cost constraints. Evaluation agents generate boundary cases from observed failures, while drift detectors identify when behaviour or input distributions move away from validated conditions. Language models help document capabilities and translate user intent into structured tasks, but hard policies and deterministic checks remain outside their control. Knowledge distillation and parameter-efficient adaptation can create smaller specialised models, while retrieval supplies changeable knowledge without forcing continual retraining. Human maintainers approve promotion, access and retirement, preserving an accountable chain from organisational need to deployed behaviour.
Model Garden
“Could organisations grow their own small, specialised intelligences?”
The beginning
Model Garden imagines artificial intelligence as an ecology of compact, governed capabilities rather than one enormous universal model. An organisation would cultivate smaller systems around specific knowledge, responsibilities and constraints: a model that understands one manufacturing process, another that checks a narrow regulatory task, another that routes a carefully defined customer request. Each intelligence has a purpose, habitat, maintainer, evaluation set and retirement condition. The metaphor matters because a garden is designed but never fully controlled. Capabilities grow through data and use, compete for resources, drift when their environment changes and sometimes need to be pruned. The objective is not maximum generality; it is a legible ecology whose behaviour an organisation can understand and govern.
General models are powerful but expensive, opaque and difficult to govern. Many business tasks require stable behaviour, private context, predictable latency and a clear accountable owner rather than the broadest possible benchmark performance. Organisations respond by adding prompts and policy documents around remote foundation models, creating the appearance of control without an operational model of what each capability knows or when it fails. Fine-tuned systems proliferate without lineage, evaluation sets become stale and teams cannot compare quality, energy, privacy and cost across alternatives. The result is a hidden monoculture: many interfaces that ultimately depend on the same upstream models, failure modes and commercial decisions.
What it could become
Model Garden would be an environment for composing datasets, compact models, retrieval systems, deterministic tools and evaluations into named governed capabilities. Every capability card declares its purpose, permitted inputs, output contract, data lineage, owner, cost profile, known limitations and conditions for retirement. A routing layer selects the smallest sufficient system for each task and falls back safely when confidence or policy boundaries are exceeded. Teams can cultivate variants in isolated plots, compare them against real failure cases and promote only those that improve the whole ecology. The interface visualises dependencies so that changing a dataset, provider or policy reveals every affected workflow before deployment.
For whom
- AI platform and engineering teams
- Regulated organisations
- Knowledge-intensive professional firms
- Product teams embedding specialised AI
Core capabilities
- Small-model training and adaptation
- Dataset lineage and consent controls
- Evaluation and behaviour monitoring
- Capability routing and orchestration
The project becomes meaningful only when a new technical possibility is translated into a clear human advantage, an experience people can understand, and a system capable of earning trust over time.
Intelligence and mathematics
Model selection is a constrained multi-objective optimisation problem. Each candidate occupies a Pareto surface defined by accuracy, calibration, latency, energy, privacy risk, maintainability and financial cost. Contextual bandits can learn routing policies while conservative exploration prevents live users from becoming uncontrolled experiments. Statistical power analysis determines whether an apparent improvement is meaningful; conformal prediction can provide task-specific uncertainty sets; and dependency graphs quantify systemic concentration risk. Ecological measures such as diversity and resilience become operational metrics: an architecture may be locally efficient yet dangerously dependent on one provider, dataset or embedding space. Mathematics makes those trade-offs inspectable rather than compressing them into a leaderboard.
How it might live
The company would sell infrastructure by governed capability and usage, beginning with organisations that already operate multiple models under meaningful privacy, reliability or regulatory constraints. The first product could be a capability registry and evaluation harness that works across providers, followed by cost-aware routing and controlled adaptation. Enterprise value would come from lower inference cost, deployment flexibility, auditable lineage and fewer duplicated experiments. Domain modules could encode specialised evaluations for legal, financial, industrial or scientific work without claiming to replace domain governance. Defensibility would emerge from evaluation history, organisation-specific failure libraries and reliable orchestration—not from owning another undifferentiated foundation model.
For me, a venture is more than an interesting technology. It needs a narrow first user, a repeated problem, a distribution path, a credible advantage and a reason to improve as more people use it. I would test those conditions before deciding whether this idea should become a company, a product, an open technology or an ongoing research programme.
Rules for making it real
- 01
Use the smallest sufficient intelligence.
- 02
No model without an owner and evaluation set.
- 03
Provenance travels with every output.
- 04
Retirement is part of the lifecycle.
From question to company
- 01Frame
Launch capability registry and evaluation harness.
- 02Prototype
Add open-model adaptation workflows.
- 03Prove
Build cost-aware routing across providers.
- 04Build
Develop regulated-industry governance modules.
What could go wrong
Serious imagination includes the possibility that an idea should change radically—or should not exist. These are the tensions the project would need to resolve:
- Fragmentation creating operational complexity.
- Small models inheriting hidden bias from weak datasets.
- Governance becoming theatre rather than engineering practice.