Daily Kickoff

Up to seven things a day for people building with AI. Fewer when there aren't seven.

  1. Chip Huyen explains how to cut inference costs without new hardware

    In a P99 conference keynote, AI Engineering author Chip Huyen breaks down inference-cost optimization into model-level techniques like quantization and distillation, and service-level ones like batching, prefill/decode separation, and prompt caching. She estimates training-to-inference compute costs run 1:10 to 1:100 over a model's life, and her tool Sniffly found 90% to 97% prompt-cache hit rates in Claude Code logs.

    The New Stack models & research

    #
  2. Firewalls run the web, just not on purpose

    Zyte's State of Web Access 2026 research found Web Application Firewalls on 92.4% of 11,100 top websites audited, with Cloudflare accounting for 34.6% of named deployments and Amazon CloudFront another 9.6%. Most run on default rulesets tuned for OWASP-style attacks rather than bot traffic, so WAF presence alone says little about how hard a site resists automated access.

    Zyte Blog

    #
  3. How software engineering is changing: an essay challenge

    The pace of change in software engineering is only accelerating, especially since January of this year. This is all to do with the industry-wide adoption of LLMs, AI tooling, and AI infrastructure.

    Pragmatic Engineer Blog

    #
  4. AI Leaders Are Calling for a Slowdown. Trump’s Team Says It’s on Them

    Sam Altman and Elon Musk this weekend backed a call for restraint, after Dario Amodei called on Washington to slow AI development. Donald Trump, meanwhile, wants the US to maintain its lead over China.

    Wired AI

    #