Software development in 2030
Testable hypotheses instead of a linear forecast
Testable hypotheses instead of a linear forecast
Testable hypotheses instead of a linear forecast
Technical Director & Fellow, large fintech
Architecture and engineering practices in fintech
AI-native development and Platform Engineering
Code of Leadership podcast, @book_cube
Test every hypothesis with four questions
Which signal can we observe now?
What must change by 2030?
What could break the scenario?
Which metric would falsify it?
But an autonomous task is not autonomous development
The question mark matters more than the arrow: stage three is unproven
Connecting an agent is easier than delegating safely
MCP connects models to context and tools
A2A describes interaction between agents
Permissions, audit and accountability stay local
Interoperability ≠ autonomy
Tool popularity is input, not outcome
No provenance for '70% use it daily'
Generated code may increase review volume
Local speed does not guarantee system gains
We need delivery, quality and outcome metrics
AlphaEvolve, WORKBank, agent economies and SE 3.0
Strong results appear where a solution can be scored automatically
DeepMind reports real results and a clear applicability condition
Data-centre, chip-design and AI-training optimisation
New matrix-multiplication algorithms
Evaluation must be objective and automated
Ambiguous product goals are not scored this way
A model being able to do a task does not mean it should be deployed
The path the authors consider most likely is also the riskiest
A human states the goal; a machine searches the solution space
This is a research programme, not a description of an era already here
Thirty goals connect sentiment and behavioural signals
The model is only one layer of an agent platform
VCS, CI/CD, observability and a service catalogue
Context, permissions and policy enforcement
Isolated execution and reproducibility
A human approves high-risk changes
Every phase needs a testable goal and an explicit decision owner
The constraint is proving quality, not generating code
Evaluators must test functionality and risk
Review cost belongs in total task cost
Rising rework falsifies the productivity promise
Safe rollback limits the blast radius
Acceptance rate helps, but does not show whether the system improved
Combine flow, experience, agent work and user outcomes
Otherwise local acceleration is easy to mistake for impact
A baseline on comparable tasks
Task success and human intervention
Total cost: compute + review + rework
Escaped defects and delivery outcome
Watch which scenario the data supports, not how much code is generated
What evidence must confirm or falsify
Intent becomes more important than typing code
Evaluators become part of development architecture
The platform defines the agent autonomy boundary
Outcomes and rework outrank generation volume
Full links are in the speaker notes
DeepMind: AlphaEvolve · Virtual Agent Economies
Stanford: Future of Work with AI Agents
SE 3.0: AI-native Software Engineering
Google Research · DORA: developer goals and metrics
polomodov.tech
All slides and links — in the Telegram channel
Alexander Polomodov, Technical Director & Fellow, large fintech
@book_cube