Harness, evals, and multi-agent mode
The first question carried over from the previous stream: do the layers around a model — reviewers, evaluators, task decomposition in markdown files — compensate for its own strength, and where does a local model stop coping? Alexander splits the question in two: decomposition and orchestration belong to the harness, evals to verifying the result. Harnesses are constantly rewritten for stronger models, and he knows of none tuned specially for a thirty-billion-parameter model on a MacBook or a consumer graphics card: building your own is slow and expensive. Evals, by contrast, pay off with any model. Savings are not guaranteed either: local inference is fine for experiments, but in multi-agent mode you run out of tokens per second. Aleksey believes a model's weakness can be compensated almost fully; the question is the price of the effort: that same day his course's Telegram group complained that a 480-billion-parameter Qwen fails on tasks the same harness clears with Opus. Alexander, quoting Boris Cherny on Claude Code, adds that a harness is built for the model that will appear in three to six months, propping up the weaknesses of one generation.
Aleksey reads the next question: how do you prove in practice that a multi-agent setup beats a single agent, and which scenarios make it worthwhile rather than wasted limits? Papers say the planner, the implementer and the reviewer work better as separate agents: one actor is biased towards confirming itself. But every agent starts up and reloads context, and on a sizeable task you can watch it re-check the re-checkers and hunt for edge cases a project without a million users will never meet. Hence the driving analogy: two hundred kilometres an hour is possible, but the chance of arriving falls, and on a dirt road even eighty is too much. The question grew out of a friend who parallelised everything through orchestration in Codex and lost track of what was happening: the agents began forwarding messages to each other. The advice is simple — do not move to several agents until the single-agent mode hits a demonstrated limit. Alexander turns on sub-agents once requirements are fixed in a spec, and gets a second opinion by asking a neighbouring model to audit the branch.
A local model and a legacy service
A question left from the earlier list: public models are banned at work and only a self-hosted Qwen is available — can it get anywhere near the top models? Alexander recalls building a local Lovable analogue on an eighty-billion-parameter model because he hated paying for foreign ones: simple pages from a prompt came out decently, but a task with a complex reference drifted away. Hence the recipe: cut tasks down to the size the model handles and explain them as you would to a junior — or push meta-information to a smart model, ask it for a plan with predictable steps, inputs, outputs and checks, and leave execution to the local one. The same split fits a data scout over a data platform his colleagues ran research over, where nobody knows in advance what is sensitive. From the chat, a listener at a European big-pharma company is looking at on-premise because of the AI Act: Aleksey notes it took effect on 2 August, requires explicit consent for using someone's data and threatens serious fines, while shadow AI on personal subscriptions flourishes inside companies.
A question from Telegram describes a mid-sized service run by three to five engineers: no documentation, requirements still unknown after eighteen months, a fragmented external context of ten to fifteen teams one handshake away, and undocumented historical workarounds. Aleksey answers that the path is almost the same as without agents. First a closed feedback loop, where reality punishes the agent with hooks and failing tests, plus a completion condition written before the start and testable from outside. Even without understanding the whole service you can describe observable behaviour: if it is a mailer, the letter must go out and carry a name rather than square brackets. Then characterization tests around the piece being changed, capturing what the system outputs today, followed by integration checks, screenshot comparisons, linters and package-dependency checks; sometimes, as in science, you have to write a tool that verifies your results. Technique does not cure organisational dysfunction: generated C4 diagrams nobody can read and nobody owns only multiply artefacts — what broke, in Aleksey's words, is not artificial intelligence but the natural kind.
Threat models, morals, and tokens
A LinkedIn question about information security: how do you avoid handing source code to third parties through a company-level LLM? Aleksey starts from the threat model — code is rarely the main secret, while vulnerabilities live along the whole chain: harness, model, tool gateway, untrusted repository, supply chain. Full isolation is possible but it is capex — your own harness, infrastructure and teams for both; the cheapest rig the participants recalled cost around a hundred thousand dollars. The compromise is a corporate gateway filtering out secrets and personal data; protection priced like a car factory is useless to a business. Alexander describes the security mindset as a staircase: the safest agent is deployed nowhere, then one without network access, then one behind a proxy, then one held to an empty allowlist. The next question: which moral principles to build into agents. Aleksey takes those that survived millennia — do no harm, do not deceive, do not slack. A strict harness catches only projections; a model has to be met at the level of meanings. Alexander answers with a book on alignment: sweeteners and contraception show how people route around evolution's own alignment.
The practical part is how to spend fewer tokens across a company. Alexander advises allocating budget to projects and tasks rather than people, costing each action and comparing it with the old process; that requires gateways with attribution and quotas. Once traces and evals exist, the gateway can route requests automatically by task type — justified by your own measurements rather than a model's public ranking. Aleksey cites an internal measurement from an engineering director he knows: pull requests reviewed by the expensive model reached production in roughly 75% of cases against 57–60% for a model twelve times cheaper; he calls the figures approximate, and it is not a public benchmark. Aleksey treats token-maxing as a false metric: what you measure is the change in production, and one interview he recalls had a company shipping around four thousand merges a day while its app rating fell. Alexander sums up: tokens are fuel, but the direction matters more, and if a team is quickly repainting buttons, the audit belongs to the backlog rather than the technology.
What to take away
- 01Evals survive a change of model; a harness does not, because it props up the weaknesses of one generation.
- 02Do not move into multi-agent mode until the single-agent one hits a demonstrated limit in speed or reliability.
- 03Prepare a legacy service for agents with a completion condition and characterization tests around the changed piece, not with generated documentation and diagrams.
- 04Price your controls and your model choice: they pay off only against a specific risk and measured quality.