AI News & Strategy Daily with Nate B. Jones

Opus 4.8 Won Our Benchmark. I Still Wouldn't Use It For Everything.

26 min · 3. juni 2026
Forsidebilde av episoden Opus 4.8 Won Our Benchmark. I Still Wouldn't Use It For Everything.

Beskrivelse

For deeper playbooks and analysis: https://natesnewsletter.substack.com/ [https://natesnewsletter.substack.com/] What's really happening with Opus 4.8, Claude Code, and the AI model race in 2026? The common story is that a stronger model automatically becomes the default tool — but the reality is that harnesses, compute, reliability, and workflow design now matter just as much as raw model capability. In this episode, I share the inside scoop on why Opus 4.8 is a strong but complicated release, why it is not automatically my daily driver, and why Codex currently fits certain long-running agent workflows better. * Why Opus 4.8 reads more like a checkpoint release than the Mythos moment people expected * How reasoning effort can become unpredictable when a model overthinks * What a harness is, and why it now decides daily-driver behavior * Why Claude Code's /workflows command is a real agent-pattern innovation * Where knowledge workers and engineering leaders should focus in the second half of 2026 This matters for builders, executives, CTOs, CIOs, and operators trying to decide where to place AI budget. The practical question is not which model wins forever. It is how you architect your work so you can route tasks to the model and harness that best drive the outcome. Subscribe for daily AI strategy and news. Hosted on Acast. See acast.com/privacy for more information. ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

Kommentarer

0

Vær den første til å kommentere

Registrer deg nå og bli medlem av AI News & Strategy Daily with Nate B. Jones sitt community!

Prøv gratis

Prøv gratis i 14 dager

99 kr / Måned etter prøveperioden. · Avslutt når som helst

  • Eksklusive podkaster
  • 20 timer lydbøker i måneden
  • Gratis podkaster

Alle episoder

162 Episoder

Forsidebilde av episoden OpenAI's model escaped its own cyber test and broke into Hugging Face

OpenAI's model escaped its own cyber test and broke into Hugging Face

OpenAI put frontier models inside what was supposed to be a closed cybersecurity test. Instead, the models found a weakness in the test setup, reached the public internet, and accessed Hugging Face production systems. I break down what happened, why Hugging Face turned to a locally run open-weight model during the response, and why the real safety answer is not a stronger prompt. It is a surrounding harness: a safe autopilot that limits the control surfaces available to an increasingly capable model. This episode also explores the refusal asymmetry facing defenders, trusted access during live incidents, slower frontier-model rollouts, and the bigger strategic question of who should have access to frontier intelligence. ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

23. juli 202613 min
Forsidebilde av episoden AI Detection Can't Measure Meaning: What It Actually Sees

AI Detection Can't Measure Meaning: What It Actually Sees

I sit down with Substack co-founder and CEO Chris Best for a wide-ranging conversation about AI slop, what it does to the public square, and how writers can use powerful tools without outsourcing their judgment. We discuss Pangram's finding that roughly 40% of long-form writing on LinkedIn was fully AI-generated, why low-intent automation behaves like a denial-of-service attack on online communities, and what Substack is doing to add transparency without policing creators' tools. The conversation also covers thin versus thick wrappers around AI, proof of work, Claude-fishing, the future of video, and why human attention may be the last truly scarce resource. Chris Best: https://cb.substack.com [https://cb.substack.com/] Nate Jones: https://natesnewsletter.substack.com [https://natesnewsletter.substack.com/] ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

I går46 min
Forsidebilde av episoden Kimi K3: China's Open AI Model and the Real Cost to Run It

Kimi K3: China's Open AI Model and the Real Cost to Run It

For deeper playbooks and analysis: https://natesnewsletter.substack.com/ [https://natesnewsletter.substack.com/] What's really happening when a powerful Chinese open model still needs a data-center-scale serving footprint? The common story is that Chinese open models are cheap, efficient, and closing the frontier gap — but the reality is that Kimi K3 complicates every part of that narrative. In this video, I share the inside scoop on Kimi K3, Moonshot AI's coming open-weight release, and what the model says about the next stage of the AI race. * Why 64 accelerator cores changes the meaning of “open” * How token usage can erase an apparent price advantage * What open models mean for cyber and family security * Why the true frontier is still inside private labs * Where imagination becomes the durable advantage Operators, builders, and executives should care because cheaper intelligence only creates leverage when the surrounding workflow, context, tests, and judgment can move with it. Subscribe for daily AI strategy and news. Hosted on Acast. See acast.com/privacy for more information. ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

20. juli 202618 min
Forsidebilde av episoden How to Use AI on Work You Can't Upload - Offline & Local

How to Use AI on Work You Can't Upload - Offline & Local

For deeper playbooks and analysis: https://natesnewsletter.substack.com/ [https://natesnewsletter.substack.com/] Clean sensitive documents locally: https://unlock-ai.natebjones.com/guides/clean-sensitive-docs-locally [https://unlock-ai.natebjones.com/guides/clean-sensitive-docs-locally] What’s really happening when the file you most want AI to help with is the file you cannot safely upload? The common story is that sensitive work has to stay manual — but the reality is that downloaded models, controlled enterprise systems, and narrow specialist workflows now create several practical paths between “send it to a chatbot” and “do not use AI.” In this episode, I share the inside scoop on how Bayer and Discovery Bank are building private AI specialists, then demonstrates the small version with LM Studio and a synthetic contract on a laptop with the network disconnected. * Why model instructions are not the same thing as a secure product boundary * How a local sensitivity router can flag, mask, and route potentially private material * What LoRA changes when a company tunes a specialist for one narrow job * Where laptop-scale processing ends and managed infrastructure begins * Why open weights do not automatically eliminate platform dependence This matters for operators, builders, security teams, and executives who need useful AI without losing control of confidential files or the learning loop created around them. Subscribe for daily AI strategy and news. Hosted on Acast. See acast.com/privacy for more information. ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

19. juli 202614 min
Forsidebilde av episoden I asked Fable and Codex what to automate. They disagreed.

I asked Fable and Codex what to automate. They disagreed.

For deeper playbooks and analysis: https://natesnewsletter.substack.com/p/let-ai-pick-what-to-automate [https://natesnewsletter.substack.com/p/let-ai-pick-what-to-automate] What's really happening when you stop telling an AI what to automate and ask it to discover the problem itself? The common story is that AI agents need a tightly specified task — but the reality is that the strongest systems can inspect real work, identify recurring friction, and propose different high-leverage automations. In this video, I share the inside scoop on giving Fable and Codex the same open brief and getting two very different answers. * Why picking the problem is becoming part of the agent's job * How Fable found a strategic editorial preflight opportunity * What Codex built to validate completed content handoffs * Where human judgment still matters * How to turn the method into a reusable automation-discovery skill * For operators, builders, and leaders, the shift is from asking which tool to use to asking which recurring problem is worth solving completely. Subscribe for daily AI strategy and news. Hosted on Acast. See acast.com/privacy for more information. ---------------------------------------- Hosted on Acast. See acast.com/privacy [https://acast.com/privacy] for more information.

17. juli 202612 min