-
Chip Huyen explains how to cut inference costs without new hardware
In a P99 conference keynote, AI Engineering author Chip Huyen breaks down inference-cost optimization into model-level techniques like quantization and distillation, and service-level ones like batching, prefill/decode separation, and prompt caching. She estimates training-to-inference compute costs run 1:10 to 1:100 over a model's life, and her tool Sniffly found 90% to 97% prompt-cache hit rates in Claude Code logs.
# -
Firewalls run the web, just not on purpose
Zyte's State of Web Access 2026 research found Web Application Firewalls on 92.4% of 11,100 top websites audited, with Cloudflare accounting for 34.6% of named deployments and Amazon CloudFront another 9.6%. Most run on default rulesets tuned for OWASP-style attacks rather than bot traffic, so WAF presence alone says little about how hard a site resists automated access.
# -
How software engineering is changing: an essay challenge
The pace of change in software engineering is only accelerating, especially since January of this year. This is all to do with the industry-wide adoption of LLMs, AI tooling, and AI infrastructure.
# -
AI Leaders Are Calling for a Slowdown. Trump’s Team Says It’s on Them
Sam Altman and Elon Musk this weekend backed a call for restraint, after Dario Amodei called on Washington to slow AI development. Donald Trump, meanwhile, wants the US to maintain its lead over China.
#