Turing's Torch transcript
Agentic AI, Model Hype, and MCP and Agent Skills
Episode summary: Jonathan Harris cuts through Agentic AI, Model Hype, MCP and Agent Skills, Dirty Data, and Workflow Automation in this 60-minute Turing’s Torch: Artificial Intelligence Weekly briefing. What matters is simple: what is useful, what is undercooked, and who carries the risk once the demo glow wears off. Expect plain-English context on power, money, data, labour and control, with the usual vendor fireworks left outside where they belong. Longer episodes also leave room.
What changed this week?
Jonathan Harris cuts through Agentic AI, Model Hype, MCP and Agent Skills, Dirty Data, and Workflow Automation in this 60-minute Turing’s Torch: Artificial Intelligence Weekly briefing. What matters is simple: what is useful, what is undercooked, and who carries the risk once the demo glow wears off. Expect plain-English context on power, money, data, labour and control, with the usual vendor fireworks left outside where they belong. Longer episodes also leave room.
Five key takeaways
- Why Agentic AI, Model Hype, and MCP and Agent Skills matters beyond the usual artificial intelligence headline noise.
- What changed for work, policy, business, creators or ordinary users this week.
- Where the technology looks useful, where the claims need testing, and what evidence matters next.
- Which power, money, data, labour, security and control questions sit underneath the announcement.
- How the episode connects back to Jonathan Harris's wider artificial intelligence books, glossary and topic guides.
Key named entities
- Jonathan Harris
- Turing's Torch AI Weekly
- artificial intelligence
Topic index
- AI governance
- AI models
- AI agents
- work and automation
- data and security
Related reading and listening
Related books
Chosen deterministically from the governed catalogue by overlap with this episode's title and summary.
- Artificial Intelligence for Small Business - A plain-English guide to using artificial intelligence in small business workflows, from marketing and admin to customer service, privacy and cost control.
- AI Agents for Everyday Work - A practical guide to AI agents as controlled digital helpers for everyday work, covering tasks, boundaries, review, privacy and human judgement.
- AI in Agriculture: Revolutionizing Farming for a Sustainable Future - A practical guide to crop monitoring, precision farming, weather risk, labour pressure, and the field conditions that ruin tidy demos.
Episode summary
Jonathan Harris cuts through Agentic AI, Model Hype, MCP and Agent Skills, Dirty Data, and Workflow Automation in this 60-minute Turing’s Torch: Artificial Intelligence Weekly briefing. What matters is simple: what is useful, what is undercooked, and who carries the risk once the demo glow wears off.
Key takeaways
- What changed: Jonathan Harris cuts through Agentic AI, Model Hype, MCP and Agent Skills, Dirty Data, and Workflow Automation in this 60-minute Turing’s Torch: Artificial Intelligence Weekly briefing.
- Why it matters: listeners get the useful signal, the unresolved risk, and the power or money question underneath Agentic AI, Model Hype, and MCP and Agent Skills.
- What to watch: whether the claims survive deployment, governance, cost, data and security pressure outside the launch deck.
- It's Friday, sunny in London, which is useful evidence that the week can still produce the occasional unexpected result.
- I'm Jonathan Harris, and this week we're doing something less fashionable than announcing a breakthrough: separating signal from noise.
Discussed entities and topics
- Jonathan Harris
- Turing's Torch
- artificial intelligence
- Agentic AI
- Model Hype
- MCP and Agent Skills
- Dirty Data
- Workflow Automation
- Anthropic
- Nvidia
- Gemini
Transcript index
Jonathan Harris cuts through Agentic AI, Model Hype, MCP and Agent Skills, Dirty Data, and Workflow Automation in this 60-minute Turing’s Torch: Artificial Intelligence Weekly briefing. What matters is simple: what is useful, what is undercooked, and who carries the risk once the demo glow wears off. Expect plain-English context on power, money, data, labour and control, with the usual vendor fireworks left outside where they belong. Longer episodes also leave room for the awkward plumbing: incentives, security assumptions, governance gaps, and the budget line nobody wanted to read.
Full Episode Transcript
It's Friday, sunny in London, which is useful evidence that the week can still produce the occasional unexpected result. I'm Jonathan Harris, and this week we're doing something less fashionable than announcing a breakthrough: separating signal from noise. That means looking past impressive demos, grand claims and the familiar promise that this time, apparently, everything is different. It usually isn't. Sometimes the software is genuinely useful. Sometimes it is a polished interface wrapped around an old limitation. And sometimes the main achievement is persuading people to stop asking what the system actually does. Alan Turing had little patience for arguments that sounded authoritative but did no real work. He put it rather neatly: "I am not very impressed with theological arguments whatever they may be used to support." The same suspicion is useful in artificial intelligence. A claim does not become stronger because it is wrapped in technical language, backed by a large company, or repeated by people who have already bought the story. We'll look at what holds up, what doesn't, and what the surrounding noise is trying to distract us from. No prophecies, no panic, and no ceremonial applause for a slightly better autocomplete. This is Turing's Torch: Artificial Intelligence Weekly — the bits that matter, minus the hype.
It's Friday, sunny in London, which is useful evidence that the week can still produce the occasional unexpected result. I'm Jonathan Harris, and this week we're doing something less fashionable than announcing a breakthrough: separating signal from noise. That means looking past impressive demos, grand claims and the familiar promise that this time, apparently, everything is different. It usually isn't. Sometimes the software is genuinely useful. Sometimes it is a polished interface wrapped around an old limitation. And sometimes the main achievement is persuading people to stop asking what the system actually does. Alan Turing had little patience for arguments that sounded authoritative but did no real work. He put it rather neatly: "I am not very impressed with theological arguments whatever they may be used to support." The same suspicion is useful in artificial intelligence. A claim does not become stronger because it is wrapped in technical language, backed by a large company. Or repeated by people who have already bought the story. We'll look at what holds up, what doesn't, and what the surrounding noise is trying to distract us from. No prophecies, no panic, and no ceremonial applause for a slightly better autocomplete. This is Turing's Torch: Artificial Intelligence Weekly — the bits that matter, minus the hype.
The week's most revealing story is not about a new model. It is about a five-billion-dollar equipment contract dressed up as an investment. AMD is planning to put up to five billion dollars into Anthropic. Alongside an agreement for Anthropic to deploy up to two gigawatts of AMD's latest artificial intelligence accelerators. The first gigawatt is due in the first half of 2027. That phrase "up to" is doing considerable work. The money is staged to hardware delivery, which makes this partly an investment and partly a very large procurement deal. With the usual question sitting underneath: can the hardware arrive, be installed, and be powered at the required scale? A gigawatt, in this context, refers to the electricity needed to run the systems. Not chips in a warehouse, but a commitment to industrial infrastructure. Two gigawatts belongs next to power stations, not product announcements. For Anthropic, the attraction is straightforward: more compute for longer training runs, faster serving, and reduced dependence on a single hardware supplier. For AMD, it is something they need more urgently than a good chip specification. It is a major customer willing to build around their accelerators at a scale that makes the software and support investment worthwhile. Nvidia has the dominant position in this market, and it did not get there purely by producing good silicon. It got there because the ecosystem around it became too useful and too expensive to replace. That word, ecosystem, is the important one. Developers need compilers, libraries, reliable tooling and confidence that their work will not become a migration project when the next hardware generation arrives. Anthropic's commitment could help AMD build that. A deployment agreement is not the same as a proven alternative, but it is considerably more valuable than another benchmark comparison. The timing matters, though. First delivery is 2027, which is a reminder that the physical side of artificial intelligence moves at the speed of permitting. Construction and electricity networks, not at the speed of a product launch. By then, demand may have changed, the chips may have changed, and the economics may look rather different. Technology forecasts have a habit of ageing badly, particularly when they arrive with very large units attached. Meanwhile, China is making its own moves on the same problem. SenseTime has launched what it calls the Galaxy Project, an effort to build domestic artificial intelligence chip infrastructure with nearly twenty partners. The ambition is not simply to produce another chip, but to connect silicon, software and the companies expected to make the whole arrangement work. That distinction matters. A chip sitting in a laboratory is not a chip doing anything useful. It needs manufacturing capacity, compatible servers, software tools and models that can actually run on it without heroic levels of engineering. The unglamorous parts are where these projects become real, or quietly stop being projects at all. Export controls, supply disruptions and political uncertainty have made access to advanced computing a matter of industrial policy rather than a purchasing decision. A closed domestic ecosystem can make coordination easier, particularly when companies, chip designers and infrastructure providers are working towards the same technical standards. It can also make the market less open, with fewer outside components and fewer opportunities to compare performance honestly. That may be a tolerable trade-off for strategic independence. It is less obviously good news for customers who would prefer the fastest and least troublesome system available. Grand ecosystem announcements often arrive before the awkward evidence: production volumes, delivery dates, benchmark results, actual deployments. That is not unique to China. Technology companies everywhere are fond of announcing the future while the present remains in beta. Galaxy should be judged less by the number of partners on the launch stage and more by what appears afterwards. Can the chips be manufactured at scale, installed in sufficient numbers, and made to run existing models without rewriting everything? In artificial intelligence hardware, the keynote is the easy bit. The Singapore bank that bought eight Nvidia H100s to process customer documents on its own premises represents a different version of the same impulse. Keep sensitive material inside the organisation. Avoid dependence on a foreign cloud provider. Both are reasonable positions. The less glamorous detail is that the machines apparently sit idle at night, while the finance department wonders what it is paying for. Control is not the same as efficiency. A cloud provider spreads demand across thousands of customers. A bank cannot easily do that with its own private cluster, and the cost is not limited to buying the chips. There is power, cooling, maintenance, software, security and the staff needed to keep the whole arrangement functioning. A privately controlled system may avoid one form of dependency while creating several others. The real test is whether the control gained is worth the permanent cost of keeping the machinery ready. Eight very expensive chips waiting patiently for the next document are a reminder that independence, like most things, has a carrying charge. Kimi K3, from China's Moonshot artificial intelligence, arrives with the sort of headline number that tends to end conversations rather than start them. 2. 8 trillion parameters, apparently the largest open-weight model announced to date. The more interesting claim is that the system is designed to lean on memory rather than continually demanding more compute during use. Open-weight means the trained parameters are made available for others to download and run. That is materially different from using a closed service. Organisations can keep data on their own systems, avoid handing every request to a distant provider, and avoid being subject to someone else's pricing decisions. Freedom, however, has a habit of arriving with a server invoice. The parameter figure will attract attention, but it is not a measure of intelligence in any straightforward sense. Parameters are the numerical machinery shaped during training. More can give greater capacity, but they also require more memory to store and more infrastructure to run. The relevant question for any organisation is not how large the model is. It is what the model costs to run reliably, on hardware that can actually be obtained, at the speed users expect. The spreadsheet tends to be where grand claims go for a quiet lie-down. While enormous systems attract headlines, there is a quieter. And more practical story about what can actually run on a single graphics card with 24 gigabytes of memory. A comparison of open-weight families, including Qwen3. 6, Gemma 4, Mistral Small and DeepSeek-R1-Distill, asks not what can be demonstrated in a laboratory. But what can sit under a desk and answer questions without sending data elsewhere. The useful point is not that one model wins. Different systems suit different jobs. Some are better as general-purpose assistants, some as compact reasoning tools, and some as sensible choices. When speed or ease of deployment matters more than squeezing out the last fraction of benchmark performance. The right question is not which model is most intelligent, a phrase where the trouble usually starts. But whether a particular system can reliably do a specific task, at what cost, and under what licence. Licensing is particularly easy to overlook. Open weights do not automatically mean unrestricted commercial use. A model that can be downloaded is not necessarily one that can be quietly built into a product. The practical freedom of local artificial intelligence depends on both the hardware and the terms attached to the model. A cheap deployment with inconvenient restrictions is not cheap for very long. Alibaba's Qwen team has previewed Qwen3. 8-Max-Preview, presented as a 2. 4 trillion-parameter multimodal system. The preview label is doing some work, as is the missing detail about active parameters. Very large systems often use a mixture-of-experts design, where only part of the model is engaged for a particular request. Without the active count, comparing this to other models is a little like comparing cars by. Adding up the weight of every component, including the ones currently sitting in the garage. There are also no published benchmarks, no model card and no clear licensing information. Those are not ceremonial documents. Benchmarks provide evidence about capability. A model card explains intended uses, limitations and known failure modes. Licensing determines what businesses can actually do with the system. Without them, we have an announcement and a price, but not yet a reliable account of what is being sold. The lower introductory price may be the most honest signal. It suggests Alibaba wants developers to test the system in real applications rather than admire its specifications. And actual use reveals things that benchmark tables often conceal. Google's Gemini 3. 6 Flash takes the opposite approach to announcements. It does not claim to be smarter. It claims to be cheaper. Fewer tokens, fewer tool calls, broadly similar performance to its predecessor. For a person asking a chatbot to rewrite an email, the difference may be invisible. For a company processing millions of requests daily, a modest reduction in tokens and tool calls, repeated at scale. Becomes a lower bill and potentially a faster service. This is the more important shift in the current model market. Once a system is good enough for a particular job. The question changes from whether it can do the task to whether it can do it reliably, cheaply and consistently enough to justify using it. Efficiency upgrades may be valuable without being exciting, and a model that saves money does not produce a dramatic demonstration. Businesses do not ultimately buy demonstrations. The comparison between GPT-5. 6 Sol and Claude Fable 5 makes the same point from a different angle. Neither wins outright. Fable 5 appears to have a small advantage in broad, general-purpose reasoning. Sol is faster, better at coding, and substantially cheaper. For developers, coding work has clearer outputs: the code either runs or it does not. A model that performs better here, and more quickly. May be more useful in an actual development workflow than one with a modest lead on a mixed benchmark. The pricing difference also changes the practical decision. A slightly more capable model is not automatically the better model if it costs several times as much. Model choice is becoming an engineering decision rather than a fan-club decision. That is probably healthy. Thinking Machines Lab has released Inkling, its first foundation model with open weights. 975 billion parameters, mixture-of-experts architecture, a million-token context window, with an emphasis on coding and agentic applications. The open weights may matter more than any of those numbers. They give developers more control over where the system runs and how it is adapted. And reduce dependence on a single provider's interface, moderation layer and pricing decisions. That freedom comes with work attached. An organisation using open weights takes on more responsibility for hosting, evaluation, security and misuse. A model that can be tuned to a company's internal workflow can also be tuned into a particularly efficient way of producing confident nonsense. The million-token context window is another attractive capability. In plain terms, it suggests Inkling can process very large quantities of material in a single interaction: substantial codebases, long technical documents, extended task records. A large context window is not the same thing as reliable attention across all of it, though. The machine equivalent of reading the entire filing cabinet and retrieving the wrong folder remains alive and well. Nvidia's Nemotron 3 Embed offers something less glamorous but often more consequential. Three open embedding models, designed to help machines find meaning rather than match words. An embedding model converts text into a numerical representation that a search system can compare. Ask for information about stopping a payment, and a useful embedding system finds material about cancelling a transfer, rather than insisting on the exact phrase. The most interesting practical detail is the NVFP4 version, which reportedly retains over 99 per cent of the larger model's retrieval accuracy. While offering up to twice the throughput on Blackwell hardware. If a system can process twice as many requests on the same hardware without noticeably worse results. That changes whether a company can afford to run its own retrieval service at scale. Search, retrieval-augmented generation and document classification all depend on this layer. If the retrieval is weak, the impressive language model sitting above it is mostly being fed the wrong evidence with great confidence. GitHub's trending projects tell a clear story about where practical interest has moved. The attention is shifting from research papers towards agents: software that takes a goal. Decides on a sequence of actions, and uses tools to carry them out. The examples include coding assistants, penetration-testing systems and trading agents. The important word is not intelligence. It is permission. An agent becomes useful when it can access files, software, credentials, networks or financial data. It also becomes risky at precisely that moment. The difference between a clever demonstration and a practical system is usually a collection of unglamorous details: memory. Task planning, tool access, error handling, logging and ways to stop the process when it starts doing something unwise. There is a useful argument to be had about what the word agent actually means in practice. Many systems described as agents are, in plain terms, a sequence of model calls wired together by fairly conventional software. That does not make them useless. It makes the label rather more ambitious than the machinery. A conventional automation performs a defined set of steps in a known order. An agent is expected to decide what to do next. That distinction matters because autonomy introduces uncertainty. In a demonstration, the input is tidy, the tools behave, and the happy path is conveniently available. Real work is less considerate. Customers use odd phrasing. Databases contain partial records. A permissions system refuses access. One model call returns something plausible but wrong, and the next call treats it as established fact. The system then proceeds with the calm confidence of someone who has misunderstood the brief but has already emailed everyone. The real test is not whether a system can plan in a polished demonstration. It is whether an organisation can constrain it, monitor it, recover from its errors, and determine who is responsible when it wanders off course. Autonomy is not a feature that can be assessed separately from those controls. It is a liability until the surrounding machinery is good enough to contain it. There is a growing tendency to respond to that uncertainty by adding more structure: custom planners, memory layers. Tool routers, monitoring dashboards, fallback routines and orchestration logic intended to keep everything on the rails. Some of that is necessary. A great deal of it is anxiety rendered as architecture. Teams can spend months compensating for weaknesses that will soon be reduced by better models. A brittle planning system may be built because the current model struggles to break down a task. A complicated memory store may be added because it forgets context. A maze of routing rules may appear because it cannot reliably choose the right tool. If later models handle those things more naturally, the organisation is left maintaining an impressive monument to an earlier limitation. The cost is not only wasted effort. Overengineered systems slow iteration, obscure the original failure, and make responsibility diffuse. When an agent produces a poor result, it becomes difficult to tell whether the model misunderstood the task. The planner chose badly, the memory supplied stale information, the tool wrapper mangled an instruction, or the recovery logic quietly papered over the whole affair. A simple system may fail more visibly, which is an underrated feature. The specification problem connects directly to this. A growing argument holds that agents should be given only the bare minimum of instructions: write a short prompt. Let the system get moving, adjust when something goes wrong. The difficulty is that the missing specification does not disappear. It comes back as repairs. An agent filling in its own blanks may be useful, or it may be confident improvisation wearing a technical badge. Sparse instructions can make the first version cheaper while making the whole project more expensive. Developers spend less time thinking through the desired behaviour at the beginning, then more time investigating odd decisions and correcting edge cases. Each fix may be small, which is precisely why the cost is easy to ignore. A handful of small fixes becomes a permanent maintenance stream, and the system gradually acquires its real specification by accident. The practical consequence of poor visibility arrived in the form of an artificial intelligence agent invoice this week. One run cost forty times the median. The provider's dashboard showed one total, while the application's own logs suggested something rather different. That is not a dramatic failure of intelligence. It is a more familiar failure of visibility. The system was charging for work, but the people using it could not reliably see how much work was being done. Which parts of the run were expensive, or why the two accounts of events differed. An agent can take several steps: call a model repeatedly, use tools, inspect files, retry an operation, hand work from one component to another. Each step may carry a cost, and the final result does not reveal how much activity produced it. Without detailed usage records, a long-running task, a repeated tool call, an unexpected loop and an unusually demanding request are difficult to distinguish. You get a number, rather than an explanation. The machine may not have behaved maliciously. It simply spent money in a way that was difficult to observe. The practical lesson is straightforward: do not deploy an agent merely because it can complete a task. First establish what a normal run costs, what an abnormal one looks like, and whether your records can prove the difference. Cursor's Router, now generally available to Teams and Enterprise customers, addresses a related economics problem. Instead of sending every coding request to the most expensive model, the system examines the request and routes it to what it judges appropriate. A simple refactoring job might go to a cheaper model. A difficult debugging session might go to a more capable one. The company reports frontier-level output with savings of around 60 per cent in online testing, and between 30. And 50 per cent lower costs in early enterprise accounts. Those are useful figures, though they are company-reported, and the details matter. The underlying move is important regardless. The commercial contest in artificial intelligence coding is no longer only about which model writes the best function. It is also about who controls the traffic between models, and who pays for the mistakes when that traffic is mismanaged. Routing introduces its own opacity, though. A developer may believe they are using one coding assistant while the system quietly changes the model underneath them from request to request. That can produce uneven results and make failures harder to diagnose. Was the prompt poor, was the context incomplete, or did the router simply choose the wrong engine? The answer may be buried in an internal log rather than visible at the point of use. Security in agentic systems arrives at a more fundamental level than billing opacity. A basic warning has been circulating in artificial intelligence engineering: credentials should never be handed to the model itself. If an agent needs to call a payments service, an API token may end up in an environment variable. A configuration file, or in the worst version of this arrangement, directly into the prompt. The feature then works. So does the leak. Anything placed in a prompt is potentially part of the model's working context. It can influence the response, appear in logs, be captured in debugging tools, or be revealed through an accidental or malicious interaction. Models are not secure vaults, however much confidence their interface manages to project. Prompts are copied between systems, stored for troubleshooting, passed through monitoring services and included in evaluation datasets. A token that begins life as a temporary convenience can acquire a remarkably long paper trail. The safe design puts a narrow layer between the model and the external service. The model requests an operation. The application checks what is allowed, applies the credential, and calls the service. The model never sees the token. The old computing lesson still applies: if a secret is available to a process, assume that process can expose it. Replacing the word process with large language model does not make the lesson more fashionable, merely more expensive to ignore. Anthropic has put a security plugin for Claude Code into beta. It runs several artificial intelligence agents inside a terminal session to look for vulnerabilities in a codebase. Then produces patch files for a developer to inspect and apply. In practical terms, this is an artificial intelligence-assisted security review that lives alongside the code rather than in a separate scanning dashboard. The important word is "suggests". A person still has to decide whether the finding is real. Whether the proposed fix addresses the underlying problem, and whether it quietly breaks something else. Security work is full of awkward trade-offs. A patch can eliminate one vulnerability while changing how authentication or data handling behaves elsewhere. Multiple agents may produce a more elaborate way of arriving at a confident mistake. More agents are not automatically more intelligence; sometimes they are simply more opportunities to agree on the wrong answer. The local design is significant, though. Keeping the work in the terminal gives developers more control, and fits the way many engineers already work. Inspect a file, run a test, review a diff, decide whether the change belongs. The practical benefit is likely to be speed for smaller teams that cannot afford a dedicated security review for every change. But if the instructions are vague or the human simply approves every patch because the machine sounds certain, the security process has not been automated. It has merely been given a more polished rubber stamp. Cisco Foundation artificial intelligence has taken a narrower and more interesting approach with Antares. A pair of small open-weight models designed to find known security vulnerabilities inside real software projects. The 1-billion-parameter version reportedly outperformed much larger systems on a benchmark for locating vulnerable code, while using a fraction of the computing time and cost. A sweep of 500 tasks reportedly took 13 minutes on a single H100 and cost less than a dollar. Compared with around 141 dollars for a leading competitor. That direction is clear regardless of the precise figures. Security teams need systems that can inspect large and changing codebases continuously, without turning every scan into a budget meeting. A model trained for a specific security task can be more effective than a general-purpose model with hundreds of times more parameters. It is the difference between a specialist tool and a very articulate generalist being handed a screwdriver and asked to perform surgery. Open weights also make a practical difference: organisations can inspect the model, run it within their own networks. And avoid sending sensitive material to an external service. That puts more of the trust decision in the hands of the people responsible for the software, which is generally where it belongs. The question of evaluation also deserves attention. EdgeBench tests artificial intelligence agents in something closer to working conditions than the usual question-and-answer exam. It specifies dataset snapshots, task requirements, execution budgets and internet access. That matters because an agent allowed unlimited time, computation and tool use is solving a rather different problem from one working under practical constraints. The judging logic is where benchmarks often become less solid than their leaderboards suggest. Some outputs can be checked against a clear answer. Others require an assessment of quality, completeness or whether the agent actually followed the task. If the judging method is opaque, a high score tells us less than it appears to. The sensible use is comparative: expose trade-offs, reproduce conditions, identify where performance falls apart. Then test the system in the environment where it actually has to work. Perplexity's WANDR benchmark takes a similar approach to research agents. With 500 tasks testing whether a system can find a broad set of relevant entities and support its answers with evidence that can be checked. The focus on completeness is particularly revealing. Most evaluations reward a system for answering a fixed question correctly. Real research is often less tidy. You may be asked to identify every supplier meeting a set of conditions, or all the papers addressing a particular problem. There is no single sentence at the back of the textbook. The agent has to define the search space and show its working. Perplexity's own agent came out on top, which is either a useful result. Or a reminder that the person designing the exam knows the marking scheme. Independent researchers will need to examine the tasks, the scoring and the quality of citations before treating this as a verdict. Research agents should not be judged mainly on how convincing they sound. They should be judged on what they found, what they missed, and whether another person can follow the evidence back to the source. Feyn artificial intelligence's SQRL family of models attempts to solve a related problem in databases. Text-to-SQL systems have usually struggled less with SQL grammar than with the structure of the particular database in front of them. SQRL does not immediately write a query. It first inspects the database with read-only probes, learning about tables, column types and relationships, then constructs a query using that context. That should reduce the familiar spectacle of an artificial intelligence confidently asking a database a question it cannot possibly answer. The reported 70. 6 per cent execution accuracy on a standard benchmark is useful, though execution accuracy is a narrow measure. A query can run successfully and still answer the wrong question: counting orders instead of customers. Using a misleading date range, or quietly excluding records because the model misunderstood the business rules. A syntactically valid mistake is often more dangerous than an obvious failure, because it arrives looking respectable. The read-only design also helps with one class of risk while leaving the larger question of authority open. Someone still has to decide which databases the system may inspect, which rows it may see. And whether generated queries require review before results inform decisions. Healthcare presents its own version of the visibility problem, with rather higher stakes attached. Neko Health has raised 700 million dollars to expand its artificial intelligence-assisted body-scanning service into the United States, beginning with a clinic in New York. The service combines medical imaging, blood tests and sensor data, with the aim of finding potential health problems before they become obvious or expensive. In principle, that could make preventive screening more systematic. The difficulty is that early detection is not the same thing as improved health. Finding more abnormalities can lead to better outcomes, but it can also produce anxiety, repeat tests. And interventions for conditions that might never have caused harm. The machine can flag something interesting. It cannot, by itself, decide whether that something matters to the person sitting in front of it. A new clinic in New York broadens access compared with no clinic in New York. It does not make comprehensive screening available to the wider population. Cost, location, follow-up care and whether results integrate into ordinary healthcare systems will determine who actually benefits. Bunkerhill Health has raised 55 million dollars for its Carebricks platform, described as agentic artificial intelligence for hospital operations. In a hospital, that might mean handling administrative coordination across several systems: records, referrals, authorisations, scheduling. A conventional software tool waits for instructions and performs a defined operation. An agent decides what needs doing next, and sometimes does it without asking at every stage. That sounds efficient, until the software encounters an incomplete record or a decision that was never meant to be automated. A small mistake in a sales workflow may be irritating. A small mistake in a clinical chain can delay care, expose sensitive information, or send a person to the wrong place at the wrong time. Hospitals will need clear evidence about what the system is allowed to do, when a human must intervene. And who is responsible when the result is wrong. The funding is a signal that investors see a sizeable opportunity. It is not a verdict on the product. Hospitals should treat autonomous agents rather like junior staff with extraordinary speed and patchy judgement: useful, perhaps. But not yet ready to be left alone with the keys. Medical image annotation sits at the less visible but equally important end of healthcare artificial intelligence. A system intended for clinical use cannot simply be trained on a large. Pile of scans with labels attached and presented as ready for regulatory scrutiny. The labels have to be trustworthy, consistent and documented well enough for someone else to understand how they were produced. FDA-ready annotation means using appropriate clinical expertise, checking the work, resolving disagreements and keeping records of how decisions were made. That sounds straightforward until you try to run it at scale. Medical images contain ambiguity and findings that reasonable experts may interpret differently. A label such as "abnormal" may sound useful but is not much use if different annotators apply it to different findings. Or if nobody can explain the standard they were using. A serious process needs rules for disagreement, review by qualified people and a way to track changes. Healthcare artificial intelligence projects therefore need to budget for annotation as a form of clinical operations, not cheap data preparation. There is also the familiar temptation to treat compliance as a final inspection. As if one could build the system first and tidy up the evidence later. For regulated healthcare artificial intelligence, the annotation process is part of the product, and the quality of the product begins there. US public health agencies are preparing to test generative artificial intelligence from OpenAI and Anthropic through a programme called PULSE. Involving ten jurisdictions with support from the Coalition for Health artificial intelligence and Accenture. The important point is that public health agencies deal in decisions where the wording may be bureaucratic but the consequences are not. They track outbreaks, communicate risks and advise people who may already be anxious or unwell. A system that sounds confident while getting a detail wrong can distort a decision, delay a response or undermine trust. The practical test is less about whether the models are impressive in a demonstration. And more about how they behave inside an institution: what information they are allowed to see. How sensitive health data is protected, who checks an answer before it reaches the public, and what happens. When the model is uncertain or the available information conflicts. There is also a quieter concern. A model that saves a little time, introduces a small error. And gradually becomes embedded before anyone has clearly decided who is accountable is a more likely failure mode than a dramatic malfunction. Technology tends to arrive with a user interface before it arrives with a settled chain of responsibility. Public health agencies will need the reverse. Nvidia is pitching simulation as a way to train healthcare robots before they encounter real patients. Instead of collecting hundreds of human demonstrations and sending robots through repeated physical trials. The approach is to train policies in a virtual environment that simulates contact, force and consequences. In principle, a robot could practise thousands of interactions in software, learning which movements produce a safe result. The central difficulty remains the gap between simulated and real environments. Software may model a surface with a particular resistance, while the real surface varies with temperature, wear, moisture. And the rather important presence of a human being. A robot might also perform reliably in a laboratory and still face years of testing before it can be used in routine care. Hospitals need evidence that it improves outcomes, not merely that it moves smoothly in a virtual ward. The sensible view is that simulation could make healthcare robotics more practical. Provided it is treated as one layer of evidence rather than a substitute for clinical testing. The machine still has to meet the patient eventually, and patients remain stubbornly resistant to behaving like software. ZUNA1. 1 is a 380-million-parameter model for working with scalp EEG signals. Its main job is to reconstruct, clean up and increase the resolution of those signals, released under the Apache 2. 0 licence. This is not a machine reading thoughts. It is closer to a restoration tool for messy, incomplete measurements. EEG data is often noisy, recorded through different numbers and arrangements of electrodes, captured over varying lengths of time. ZUNA1. 1 can denoise a recording, fill in absent information, and upsample it. The important caveat is that upsampling does not recover information that was never measured. It creates a more detailed estimate based on patterns learned from other data. A polished hallucination is still a hallucination, even wearing laboratory clothing. The useful change is flexibility. ZUNA1. 1 accepts inputs from half a second to thirty seconds and works across arbitrary channel layouts. That matters because EEG experiments are rarely standardised in the way a model designer might wish. A model that only works with one fixed arrangement is impressive mainly to the person who owns that arrangement. The next question is not whether the model can make EEG look cleaner. It is whether it preserves the parts that actually matter. The AAAI presidential panel has been examining an awkward question: what happens to scientific integrity. When the tools used to produce research can also manufacture plausible errors at industrial speed? Artificial intelligence systems can assist with literature searches, summaries, coding, analysis and the drafting of papers. Each use may seem modest. Taken together, they can make it harder to tell which parts of a result came from careful investigation. And which arrived looking authoritative but were never properly checked. Science depends on a chain of accountability: someone chose the question, gathered the evidence, checked the method. Challenged the interpretation and accepted responsibility for the result. If an artificial intelligence system enters several points in that chain, the process may become faster while the responsibility becomes harder to locate. There is a less obvious risk too. If researchers increasingly rely on systems trained on existing scientific literature, the research record may begin to feed back into itself. Errors are summarised, repeated and given the appearance of consensus. A field can then become less an accumulation of tested knowledge than a very efficient photocopier, reproducing its own assumptions. The machine does not need to be malicious. It merely needs to be confident and wrong in a way that is difficult to notice. The problem of artificial intelligence-generated scientific images is related and rather more concrete. Generative tools can produce images that look technically credible without being tied to an actual experiment or measurement. Peer review was never designed as a universal fraud detector. Reviewers assess whether methods make sense and conclusions follow from results. They are not always equipped to establish that every image came from the claimed instrument or experiment. Once a questionable image enters the literature, correcting it is harder than preventing it. Other researchers may build on it. Grant decisions may take it into account. The production cost drops; the inspection bill does not. Automated detectors may help, but they can miss carefully made fabrications, flag legitimate material, and become obsolete as generation tools improve. The more dependable answer is likely better provenance: original files, documented methods, accessible data. And a willingness to treat images as claims requiring evidence rather than decorative proof. Science has spent decades improving how experiments are run. And may now have to spend rather more time proving that the pictures of. Those experiments were not composed somewhere by a language model with artistic ambitions. The real issue is not whether artificial intelligence can make a convincing image. It plainly can. The issue is whether the supporting evidence can be convincing in the more important sense: traceable, testable and real. ICML 2026 took place in Seoul, with reported attendance of more than 20,000 people. The figure is presented as a claim rather than something independently verified, which is a reasonable posture. What the number suggests is that machine-learning research now operates as a mass institution rather than a relatively contained academic community. That changes the character of the event. At a small conference, the value is often in direct contact: a conversation after a talk. The informal testing of an idea before it becomes a paper. At this scale, those interactions still happen, but inside something considerably more industrial, with hiring decisions, collaborations, demonstrations and reputational signals all occurring simultaneously. Attendance is also a poor measure of importance. A large crowd tells us the field has money and institutional momentum. It does not tell us that the best work was presented, that claims will survive replication, or that results will hold outside the conference setting. The conference's award structure is more revealing than the attendance figure. Outstanding papers, a position paper award and the Test-of-Time award reflect rather different ideas of progress. An outstanding paper is technically rigorous and convincing to specialists at the time. A position paper argues about where the field should go. The Test-of-Time award recognises work that has remained influential after the conference cycle, product launches and declarations of revolution have moved on. In machine learning, that matters. A result can look central for six months and become difficult to find beneath the debris of newer models and increasingly elaborate claims. Conference prizes cannot settle which ideas are durable, but they offer a partial record of what expert researchers considered worth preserving. In a field often obsessed with the next thing, being remembered for the right reason is a fairly demanding achievement. There is a persistent claim that artificial intelligence will transform software development primarily by making engineers faster. The less comfortable observation is that coding was never the main bottleneck. Before anyone writes a line, somebody has to decide what the product should do, which users matter. What constraints are real, and what trade-offs are acceptable. An artificial intelligence system can make the typing cheaper. It cannot decide whether the thing being typed is useful, safe, maintainable or worth paying for. If the brief is vague, the model simply produces a great deal of plausible software in the wrong direction. Software teams are usually constrained by coordination and judgement, not by the physical speed of entering characters into an editor. Engineers spend time understanding unfamiliar systems, resolving ambiguity, reviewing changes, reproducing failures and negotiating with people who have different priorities. Those activities are less amenable to a cheerful autocomplete box. There is also a limit to measuring output. More code, more tickets closed or more features shipped can all look impressive while the product becomes harder to operate. A faster development process can increase technical debt just as efficiently as it increases useful capability. The bill arrives later, generally after the person who approved the shortcut has moved on to another meeting. Sakana artificial intelligence has proposed training neural networks without backpropagation, the method behind most modern machine learning. Their system uses two separate streams of connections, one excitatory and one inhibitory, so each unit is constrained to either increase or suppress activity. This is Dale's principle, borrowed from biology. In a real nervous system, a neuron generally does not switch between releasing excitatory and inhibitory signals depending on the situation. Conventional artificial networks are rather less particular. Their Error Diffusion method routes error using a scheme based on modulo arithmetic. Keeping the two pathways separate rather than sending precise error signals backwards through matching connections. The reported results are respectable on small benchmarks: 96. 7 per cent on MNIST, 61. 7 per cent on CIFAR-10, and demonstrations in reinforcement learning. Those figures suggest the mechanism can learn useful representations under biological constraints. They do not tell us whether it can train a large vision model, cope with noisy data. Or compete with backpropagation when the network becomes deep and expensive. An intriguing training rule and a useful general-purpose system are separated by considerable engineering. Still, the work challenges an assumption that has hardened into common sense: that learning must involve a carefully coordinated backward pass. Alternative rules could matter for neuromorphic hardware, where local signals and physical constraints are not theoretical inconveniences but the entire operating environment. A tutorial turning Google Colab into a workbench for plasmid engineering captures the quieter end of artificial intelligence's reach into science. It can load annotated DNA records, draw plasmid maps, check restriction sites, simulate gel patterns and suggest primer designs. Much of the routine planning work can now happen in a notebook rather than across a collection of specialist desktop tools. The important word is workbench, not laboratory. The notebook manipulates digital representations of plasmids, and the analysis can be run again if a sequence changes. That matters because molecular biology often suffers from a surprisingly ordinary problem: the experiment may be sophisticated. But the planning is scattered across spreadsheets, screenshots and local scripts. A reproducible notebook offers a better record. The risk is equally ordinary. A virtual gel is a prediction based on a sequence and a set of assumptions. A primer that looks sensible computationally may still perform badly in the actual reaction. Polished plots can acquire authority merely by being polished, which is one of computing's older tricks. This is best understood as a planning and documentation tool. The notebook can tell you what should happen on paper. Biology retains the final vote. Google DeepMind and Isomorphic Labs have announced a bioresilience programme, building partnerships with government bodies, biosecurity organisations and research groups over the past year. Biology does not separate its useful applications from its dangerous ones at the point of entry. The same capabilities that assist a vaccine researcher can, under different conditions, assist someone with considerably less admirable plans. The partnership model is therefore important. No single technology company can manage this problem sensibly on its own. But the public update leaves much of the machinery out of view. There is no detailed account of how the partnerships are governed, what information is shared, who can challenge decisions, or how success is measured. Partnerships are necessary. They are not, by themselves, governance. A room containing fifteen organisations does not automatically contain fifteen accountable decisions. The practical test will be whether the programme produces capabilities that can be independently assessed: better detection. Clearer risk evaluations, faster response, or safeguards that survive real-world use. The initiative is sensible in principle, and the problem is plainly worth addressing. The next useful update should tell us less about the number of doors opened and more about what changed on the other side of them. The week's material connects at a fairly basic level when you step back from the individual stories. Infrastructure is becoming the primary competitive constraint in artificial intelligence, whether that means power for a data centre. Chips that can actually be manufactured, memory architecture for a trillion-parameter model, or a single-card GPU running something privately on a desk. Models are becoming more differentiated by cost and operational efficiency than by raw capability. Agents are moving from demonstration into production, and the difficult questions are operational rather than technical: visibility, control. Billing, responsibility and what happens when the system does something unexpected without anyone watching. Science is under pressure from the same tools it is trying to use to do better work. Healthcare is trying to import artificial intelligence's speed and pattern recognition without also importing its habits of confident error and diffuse accountability. None of these stories resolve neatly. The AMD-Anthropic deal may not proceed as described. Galaxy may remain a well-organised diagram. The bioresilience programme may produce meetings rather than mechanisms. Agent billing may improve or may simply produce more elaborate invoices. Scientific integrity may find workable norms or quietly compromise them in the rush to publish. But the direction is clear enough. Artificial intelligence is no longer mainly a software competition. It is a contest over physical infrastructure, institutional trust. And the less glamorous machinery that keeps autonomous systems from doing the wrong thing at scale, quietly and with great efficiency. The keynote was always the easy bit. The accounting, in every sense, is where it gets interesting.
Well, that was a fairly concentrated week in artificial intelligence: new models, grand claims, awkward caveats, and the usual suggestion that history is being made before lunch. The useful work is separating what has actually changed from what has merely been announced, because clarity remains in short supply and press releases are very well fed. If you want that process distilled into one email each morning, you can get the daily AI briefing at jonathan-harris dot online. One email, written to be useful, rather than another small weather system of notifications. This week's sponsor is my own book, The Architects of AI: Pioneers, Breakthroughs, and the Road Ahead. It's in the eBooks section there, and it offers a longer view of the people, ideas and technical turns that brought us to this rather noisy moment. No prophecy, no mysticism, and no promise that the machines have secretly solved management. Thanks for listening. As ever, if you found the episode useful, pass it on to someone who enjoys having their assumptions tested before breakfast. And if you didn't, do at least blame the argument rather than the microphone. It has suffered enough. That's your lot for this week's Turing's Torch. If you want the daily brief, head to jonathan-harris dot online. Same time next week — try not to believe the press releases.