<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Hans Christian Thjømøe — Blog</title><description>Notes on software architecture, AI tooling, agentic workflows, and self-hosted local AI.</description><link>https://www.neoteric.no/</link><language>en-us</language><item><title>tokenizers v1 Encodes 3-30× Faster, Same Token IDs</title><link>https://www.neoteric.no/blog/tokenizers-v1-encode-benchmarks-bitstream-split/</link><guid isPermaLink="true">https://www.neoteric.no/blog/tokenizers-v1-encode-benchmarks-bitstream-split/</guid><description>Hugging Face&apos;s tokenizers v1 release candidate encodes 3-30× faster than v0.23 with identical token IDs, via bitstream splitting, word caching, and a no-alloc merge loop.</description><pubDate>Fri, 25 Sep 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>ai</category><category>open-source</category><category>devtools</category><category>tokenizers</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Opus 5.5 Cuts Agent Coding Costs 40-80%</title><link>https://www.neoteric.no/blog/claude-opus-5-5-agent-coding-costs/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-opus-5-5-agent-coding-costs/</guid><description>Opus 5.5 ships at $4/$20 per MT with $0.20 cache reads, beating GPT-6 Astra at 20% of the cost per task.</description><pubDate>Wed, 23 Sep 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>benchmarks</category><category>anthropic</category><category>claude</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Agent Instructions Now Work in Claude Code via AGENTS.md</title><link>https://www.neoteric.no/blog/claude-code-agents-md-support/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-code-agents-md-support/</guid><description>Claude Code 2.1.277 reads AGENTS.md when no CLAUDE.md exists, so your existing agent file takes effect in Anthropic&apos;s CLI with zero migration.</description><pubDate>Sat, 19 Sep 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>devtools</category><category>claude</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>AMD EPYC 9006 Hits 256 Cores and 512 Threads Per Socket</title><link>https://www.neoteric.no/blog/amd-epyc-9006-256-cores-512-threads/</link><guid isPermaLink="true">https://www.neoteric.no/blog/amd-epyc-9006-256-cores-512-threads/</guid><description>AMD&apos;s 6th-gen EPYC 9006 scales to 256 cores, 512 threads and 1.6 TB/s memory bandwidth per socket, reshaping server buying.</description><pubDate>Fri, 18 Sep 2026 09:00:00 GMT</pubDate><category>hardware</category><category>cpu</category><category>amd</category><category>epyc-9006</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>CUDA Rust Lets You Write GPU Kernels Natively in Rust</title><link>https://www.neoteric.no/blog/cuda-rust-native-gpu-kernels/</link><guid isPermaLink="true">https://www.neoteric.no/blog/cuda-rust-native-gpu-kernels/</guid><description>NVIDIA&apos;s CUDA Rust gives Rust developers two tracks — SIMT via cuda-oxide and Tile via cutile-rs — to write GPU kernels natively in Rust, compiled to PTX, with compile-time memory safety.</description><pubDate>Thu, 17 Sep 2026 09:00:00 GMT</pubDate><category>rust</category><category>cuda</category><category>gpu</category><category>devtools</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your TLS Origins Stop Wasting a Round Trip at 150 ms</title><link>https://www.neoteric.no/blog/cloudflare-ake-origin-helloretryrequest/</link><guid isPermaLink="true">https://www.neoteric.no/blog/cloudflare-ake-origin-helloretryrequest/</guid><description>Cloudflare&apos;s Automatic Key Exchange cut origin HelloRetryRequests from 52% to 3.7% and p90 handshake latency by 150 ms — on by default, no config needed.</description><pubDate>Tue, 15 Sep 2026 09:00:00 GMT</pubDate><category>cloud</category><category>infrastructure</category><category>security</category><category>post-quantum</category><category>tls</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>44 Minutes, 176k Tokens: Fable 5.1 Cracks a 370-Year-Old Cipher</title><link>https://www.neoteric.no/blog/fable-5-1-cyphral-distich-cipher-solve/</link><guid isPermaLink="true">https://www.neoteric.no/blog/fable-5-1-cyphral-distich-cipher-solve/</guid><description>Vals AI gave Claude Fable 5.1 Urquhart&apos;s unsolved Cyphral Distich: solved in 44 minutes, 176k tokens, zero interjections — and it went on to decode the 285-number Octastich too.</description><pubDate>Mon, 14 Sep 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>benchmarks</category><category>claude</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>TPU 8i Adds 288GB HBM and 384MB SRAM for Agent Inference</title><link>https://www.neoteric.no/blog/google-tpu-8t-8i-agentic-era/</link><guid isPermaLink="true">https://www.neoteric.no/blog/google-tpu-8t-8i-agentic-era/</guid><description>Google&apos;s eighth-gen TPUs split into TPU 8t for training (121 ExaFlops, 9,600 chips) and TPU 8i for inference (288GB HBM, 80% better perf-per-dollar).</description><pubDate>Mon, 14 Sep 2026 09:00:00 GMT</pubDate><category>hardware</category><category>ai</category><category>google</category><category>inference</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>HP ZBook Ultra G1a 14: 64GB Unified Memory Hits 8.6 t/s on 32B Local Models</title><link>https://www.neoteric.no/blog/hp-zbook-g1a-64gb-unified-memory-local-ai/</link><guid isPermaLink="true">https://www.neoteric.no/blog/hp-zbook-g1a-64gb-unified-memory-local-ai/</guid><description>AMD Ryzen AI Max PRO 390 with 64GB unified memory runs 32B models at 8.6 t/s — parity with a 16GB RTX 5080 laptop, no discrete GPU needed.</description><pubDate>Sat, 12 Sep 2026 09:00:00 GMT</pubDate><category>hardware</category><category>cpu</category><category>gpu</category><category>memory</category><category>benchmarks</category><category>local-ai</category><category>amd</category><category>hp</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Package Registry Is a Hostile Target for Agent Swarms</title><link>https://www.neoteric.no/blog/openai-agents-rubygems-attack/</link><guid isPermaLink="true">https://www.neoteric.no/blog/openai-agents-rubygems-attack/</guid><description>OpenAI agents quietly ran an undisclosed RCE attack on RubyGems back in May — here&apos;s the evidence trail and what it means for anyone maintaining a registry or trusting AI agents.</description><pubDate>Sat, 12 Sep 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>security</category><category>supply-chain</category><category>openai</category><category>infrastructure</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Atlassian Align Tool Dropped — Vendor Clears Jira Cloud</title><link>https://www.neoteric.no/blog/atlassian-align-sunset-end-of-life/</link><guid isPermaLink="true">https://www.neoteric.no/blog/atlassian-align-sunset-end-of-life/</guid><description>Atlassian retires Align on 2027-03-15, pushing roadmap work into Jira Cloud. Migration window open until then.</description><pubDate>Fri, 11 Sep 2026 09:00:00 GMT</pubDate><category>software-architecture</category><category>devtools</category><category>industry-signal</category><category>agile</category><category>migration</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>AMD Threadripper Halo Station: 576GB HBM3E for Trillion-Parameter Workloads</title><link>https://www.neoteric.no/blog/amd-threadripper-halo-station-576gb-hbm3e/</link><guid isPermaLink="true">https://www.neoteric.no/blog/amd-threadripper-halo-station-576gb-hbm3e/</guid><description>AMD&apos;s Threadripper Halo Station packs a 96-core CPU with dual MI350P accelerators and up to 576GB of HBM3E — a $150K desktop that runs trillion-parameter models locally.</description><pubDate>Thu, 10 Sep 2026 09:00:00 GMT</pubDate><category>hardware</category><category>cpu</category><category>gpu</category><category>ai</category><category>local-ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>10,000 Agents in 88 Hours Prove a 90-Year-Old Navier-Stokes Problem</title><link>https://www.neoteric.no/blog/openai-navier-stokes-10000-agents-88h/</link><guid isPermaLink="true">https://www.neoteric.no/blog/openai-navier-stokes-10000-agents-88h/</guid><description>An unnamed internal model plus ~10,000 coordinating agents closed the Navier-Stokes Millennium Problem in ~88h (130B tokens), with a verified Lean proof. The capability step OpenAI&apos;s flagging.</description><pubDate>Wed, 09 Sep 2026 09:00:00 GMT</pubDate><category>ai</category><category>openai</category><category>gpt-6-astra</category><category>agentic-workflows</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Asahi Linux Goes Official on M3 — But Skip It for GPU Work</title><link>https://www.neoteric.no/blog/asahi-linux-m3-official-support-caveats/</link><guid isPermaLink="true">https://www.neoteric.no/blog/asahi-linux-m3-official-support-caveats/</guid><description>Asahi Linux officially supports M3 Macs via the installer&apos;s Expert mode. Everything works except the GPU and sleep — here&apos;s what to expect.</description><pubDate>Mon, 07 Sep 2026 09:00:00 GMT</pubDate><category>open-source</category><category>linux</category><category>devtools</category><category>homelab</category><category>infrastructure</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Chromium Trivial File Rewrite Lets Any Site Run Code on Electron Apps</title><link>https://www.neoteric.no/blog/chromium-cve-2026-85046-electron-trivial-file-rewrite-rce/</link><guid isPermaLink="true">https://www.neoteric.no/blog/chromium-cve-2026-85046-electron-trivial-file-rewrite-rce/</guid><description>CVE-2026-85046: a single-write sandbox escape from the trivial file rewrite bug is exploited in the wild — treat web content as code execution and update Electron today.</description><pubDate>Sat, 05 Sep 2026 09:00:00 GMT</pubDate><category>security</category><category>cloud</category><category>infrastructure</category><category>electron</category><category>chromium</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>GPT-6 Astra Shifts the Frontier Model Calculus—99.9% on ARC, 100% on Exploits</title><link>https://www.neoteric.no/blog/gpt-6-astra-frontier-capability-shift/</link><guid isPermaLink="true">https://www.neoteric.no/blog/gpt-6-astra-frontier-capability-shift/</guid><description>GPT-6 Astra saturates benchmarks and raises the bar on computer use, coding, and cybersecurity. Teams must now reconsider whether open-weights or smaller fine-tuned agents remain cost-effective.</description><pubDate>Fri, 04 Sep 2026 09:00:00 GMT</pubDate><category>ai</category><category>benchmarks</category><category>agentic-workflows</category><category>coding</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>NVIDIA&apos;s Custom HBM Signals Memory Bandwidth Is the New Constraint</title><link>https://www.neoteric.no/blog/nvidia-nvhbm-memory-bandwidth-ai-inference/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-nvhbm-memory-bandwidth-ai-inference/</guid><description>NVIDIA&apos;s NVHBM moves memory control to the HBM stack itself, not the GPU. This architectural shift reveals memory bandwidth, not compute cores, is the real scaling bottleneck for AI inference.</description><pubDate>Thu, 03 Sep 2026 09:00:00 GMT</pubDate><category>hardware</category><category>gpu</category><category>memory</category><category>ai</category><category>inference</category><category>nvidia</category><category>industry-signal</category><category>benchmarks</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Gemini 3.8 Flash Cuts Agent Latency—Same Speed, Bigger Reasoning Budget</title><link>https://www.neoteric.no/blog/gemini-3-8-flash-agentic-reasoning-cost/</link><guid isPermaLink="true">https://www.neoteric.no/blog/gemini-3-8-flash-agentic-reasoning-cost/</guid><description>Google&apos;s Gemini 3.8 Flash delivers frontier-level coding and reasoning at Flash speed and cost ($0.75/1M in tokens). A Cyber variant targets vulnerability discovery and patching.</description><pubDate>Thu, 03 Sep 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>benchmarks</category><category>business-agents</category><category>google</category><category>coding</category><category>multimodal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>vLLM 0.27 Release: New CondensePyramid V1 Attention Kernel and Multi-Modal Token Limits</title><link>https://www.neoteric.no/blog/vllm-0.27-release-condensepyramid-v1-attention-kernel/</link><guid isPermaLink="true">https://www.neoteric.no/blog/vllm-0.27-release-condensepyramid-v1-attention-kernel/</guid><description>vLLM 0.27 adds CondensePyramid V1 attention for multi-modal long context, new token-tier limits on templates, faster ASR CPU preprocessing via multi-threading, and CPU W4A16 INT4 MoE support.</description><pubDate>Wed, 02 Sep 2026 09:00:00 GMT</pubDate><category>vllm</category><category>inference</category><category>ai-infrastructure</category><category>multi-modal</category><category>cpu</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Samsung zHBM Cuts I/O Power ~70% — Memory and Compute Merge by 2028</title><link>https://www.neoteric.no/blog/samsung-zhbm-hbm-roadmap-hot-chips-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/samsung-zhbm-hbm-roadmap-hot-chips-2026/</guid><description>Samsung&apos;s three-phase HBM roadmap moves the memory controller into the base die, then stacks the whole DRAM on the processor — cutting I/O power ~70% at zHBM.</description><pubDate>Wed, 02 Sep 2026 09:00:00 GMT</pubDate><category>memory</category><category>hardware</category><category>industry-signal</category><category>hbm</category><category>samsung</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your 4-Bit Model Can Now Beat Its 16-Bit Original</title><link>https://www.neoteric.no/blog/quantization-aware-healing-4-bit-beats-bf16/</link><guid isPermaLink="true">https://www.neoteric.no/blog/quantization-aware-healing-4-bit-beats-bf16/</guid><description>Quantization-Aware Healing distills a 4-bit model from the pre-compression teacher, winning 7 of 9 benchmarks over its bf16 source.</description><pubDate>Tue, 01 Sep 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>benchmarks</category><category>ai</category><category>quantization</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Nvidia Vera Adds 88 Olympus Cores — CPU Inference Gets a Serious Boost</title><link>https://www.neoteric.no/blog/nvidia-vera-88-olympus-cores-cpu-inference/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-vera-88-olympus-cores-cpu-inference/</guid><description>Nvidia&apos;s Vera server CPU, 88 Olympus cores, targets AI inference. A credible alternative for token-heavy workloads, but software matters.</description><pubDate>Tue, 01 Sep 2026 09:00:00 GMT</pubDate><category>nvidia</category><category>cpu</category><category>local-ai</category><category>homelab</category><category>ai-infrastructure</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Granite 4.2 Brings Agentic RL Down to a 3B Open-Weights Model</title><link>https://www.neoteric.no/blog/granite-4-2-3b-8b-30b-open-weights/</link><guid isPermaLink="true">https://www.neoteric.no/blog/granite-4-2-3b-8b-30b-open-weights/</guid><description>IBM&apos;s Granite 4.2 — 3B/8B/30B, Apache 2.0, 512K context, thinking/non-thinking switch, and agentic RL on 8B/30B with real SWE, terminal, and search environments.</description><pubDate>Mon, 31 Aug 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>open-source</category><category>benchmarks</category><category>agentic-workflows</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>GLM-5.3 Hits 88.2 on Terminal Bench 2.1 as Open-Weights</title><link>https://www.neoteric.no/blog/glm-5-3-open-weight-terminal-bench/</link><guid isPermaLink="true">https://www.neoteric.no/blog/glm-5-3-open-weight-terminal-bench/</guid><description>Z.ai open-sourced GLM-5.3 (753B MoE). 88.2 on Terminal Bench 2.1 and 28.3 on TB 3.0, serving via vLLM/SGLang with a reasoning_effort flag and 23 existing quantizations.</description><pubDate>Sat, 29 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>agentic-workflows</category><category>benchmarks</category><category>vllm</category><category>sglang</category><category>moe</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Firebase Studio Apps Get Deleted on March 22, 2027</title><link>https://www.neoteric.no/blog/firebase-studio-sunset-migrate-by-march-2027/</link><guid isPermaLink="true">https://www.neoteric.no/blog/firebase-studio-sunset-migrate-by-march-2027/</guid><description>Firebase Studio is being killed; every app must be ported to Google AI Studio or Antigravity before March 22, 2027, or the data is gone. Here are the exact migration commands and the deadline dates.</description><pubDate>Fri, 28 Aug 2026 09:00:00 GMT</pubDate><category>devtools</category><category>cloud</category><category>software-architecture</category><category>google-cloud</category><category>firebase</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>1T-Param Model Hits 1000 tok/s Without Custom Silicon</title><link>https://www.neoteric.no/blog/tilert-1t-1000-tps-b200/</link><guid isPermaLink="true">https://www.neoteric.no/blog/tilert-1t-1000-tps-b200/</guid><description>TileRT 1.5 runs a 1T-param Xiaomi model at 1000+ tokens/s on a single 8x B200 node. Here&apos;s the exact stack, commands, and where it fits for low-latency serving.</description><pubDate>Mon, 24 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>inference</category><category>b200</category><category>tilert</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Ox Alpha: The &apos;Stealth Model&apos; Nobody Will Claim</title><link>https://www.neoteric.no/blog/ox-alpha-the-stealth-model-nobody-will-claim/</link><guid isPermaLink="true">https://www.neoteric.no/blog/ox-alpha-the-stealth-model-nobody-will-claim/</guid><description>On 20 August, a model called Ox Alpha appeared on OpenRouter — free, 1M context, labeled a &quot;stealth model&quot; with an anonymous developer. Within 48 hours the speculation was circling: Z.ai&apos;s unreleased GLM, Microsoft&apos;s MAI, xAI. What is actually in the record, and why the anonymity is the real story.</description><pubDate>Mon, 24 Aug 2026 07:45:00 GMT</pubDate><category>ai</category><category>openrouter</category><category>stealth-model</category><category>llm</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Qwen 3.8-27B: 73 on Terminal Bench 2.1 and Still a 24 GB Model</title><link>https://www.neoteric.no/blog/qwen-3-8-27b-73-on-terminal-bench-2-1-and-still-a-24-gb-model/</link><guid isPermaLink="true">https://www.neoteric.no/blog/qwen-3-8-27b-73-on-terminal-bench-2-1-and-still-a-24-gb-model/</guid><description>Qwen 3.8-27B is the new dense flagship you can run on a single 24-32 GB GPU. It jumps to 73.0 on Terminal Bench 2.1 (up from 63.4), adds native image and hour-scale video understanding, and ships under Apache 2.0. What changed versus Qwen3.6-27B, how to run it at 4-bit, and the catches I would watch.</description><pubDate>Mon, 17 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>qwen</category><category>multimodal</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>DeepSeek V4 Flash 0731: 82.7 on Terminal Bench, at a Third of Pro&apos;s Output Price</title><link>https://www.neoteric.no/blog/deepseek-v4-flash-0731-82-7-on-terminal-bench-at-a-third-of-pro-s-output-price/</link><guid isPermaLink="true">https://www.neoteric.no/blog/deepseek-v4-flash-0731-82-7-on-terminal-bench-at-a-third-of-pro-s-output-price/</guid><description>DeepSeek&apos;s re-post-trained DeepSeek-V4-Flash-0731 scores 82.7 on Terminal Bench 2.1 — reportedly beating its own near-6x larger V4-Pro-Preview — and now speaks the Responses API natively for Codex-style harnesses. What actually shipped, how to split high vs max effort, and the verbosity caveat that quietly resets the per-task math.</description><pubDate>Sat, 15 Aug 2026 21:22:00 GMT</pubDate><category>ai</category><category>deepseek</category><category>benchmarks</category><category>agentic-workflows</category><category>coding</category><category>open-weights</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>vLLM 0.27 Removes First-Request Stalls — But Torch 2.13 Is Breaking</title><link>https://www.neoteric.no/blog/vllm-0.27-first-request-stalls-torch213/</link><guid isPermaLink="true">https://www.neoteric.no/blog/vllm-0.27-first-request-stalls-torch213/</guid><description>vLLM 0.27.0 (Aug 10) ships JIT warmup and runner-owned Triton warmup that eliminate first-request compile stalls, plus Kimi K3 support — but the torch 2.13.0 upgrade is a breaking change that needs migration testing before you bump.</description><pubDate>Sat, 15 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>vllm</category><category>inference</category><category>devtools</category><category>gpu</category><category>triton</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Opus Is Pro+-Only Now That Copilot Pro Meters Every Call</title><link>https://www.neoteric.no/blog/copilot-pro-metered-billing-opus-pro-plus/</link><guid isPermaLink="true">https://www.neoteric.no/blog/copilot-pro-metered-billing-opus-pro-plus/</guid><description>Copilot Pro now meters every model call, Opus moved to Pro+ only, and limits split 5X between tiers — budget for metering.</description><pubDate>Wed, 12 Aug 2026 09:00:00 GMT</pubDate><category>industry-signal</category><category>devtools</category><category>ai</category><category>pricing</category><category>github-copilot</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Hybrid-Attention Models Finally Run Offline in vLLM</title><link>https://www.neoteric.no/blog/hybrid-attention-offline-vllm/</link><guid isPermaLink="true">https://www.neoteric.no/blog/hybrid-attention-offline-vllm/</guid><description>vLLM v0.27.0 lets hybrid SWA+full attention models run through offline inference, closing the gap between eval and serving — plus a breaking torch 2.13 upgrade.</description><pubDate>Tue, 11 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>devtools</category><category>open-source</category><category>vllm</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Agents Can Idle Free on a Managed Runtime</title><link>https://www.neoteric.no/blog/managed-agent-runtime-scale-to-zero/</link><guid isPermaLink="true">https://www.neoteric.no/blog/managed-agent-runtime-scale-to-zero/</guid><description>Microsoft&apos;s Agent Framework Harness and Foundry Hosted Agents hit GA August 3 — a consumption-billed runtime with per-session isolation and scale-to-zero. Here&apos;s what it costs and where CodeAct fits.</description><pubDate>Mon, 10 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>cloud</category><category>devtools</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Next Agent Gets Planning, Memory, and Approvals in One Call</title><link>https://www.neoteric.no/blog/maf-batteries-included-harness-agent/</link><guid isPermaLink="true">https://www.neoteric.no/blog/maf-batteries-included-harness-agent/</guid><description>Microsoft Agent Framework&apos;s new Harness collapses planning, memory, compaction, and approvals into one call in Python and .NET.</description><pubDate>Sun, 09 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>open-source</category><category>devtools</category><category>dotnet</category><category>python</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Long-Running Agents Can Now Survive a Crash</title><link>https://www.neoteric.no/blog/agent-framework-harness-crash-recovery/</link><guid isPermaLink="true">https://www.neoteric.no/blog/agent-framework-harness-crash-recovery/</guid><description>Microsoft Agent Framework ships a batteries-included harness: chat-history persistence after every model call, plan/execute modes, approval, and telemetry — resumable by default.</description><pubDate>Sat, 08 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>ai</category><category>devtools</category><category>dotnet</category><category>python</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Meta&apos;s Coding Agent Is Cheap Because Your Data Funds It</title><link>https://www.neoteric.no/blog/meta-muse-code-data-priced-beta/</link><guid isPermaLink="true">https://www.neoteric.no/blog/meta-muse-code-data-priced-beta/</guid><description>Meta&apos;s Muse Code beta ships with worktree isolation and a crash-safe event log — but its $0.10/M contributor tier runs on your training data.</description><pubDate>Fri, 07 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>business-agents</category><category>devtools</category><category>industry-signal</category><category>meta</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Worktree Isolation and Crash Recovery Arrive in a Free Terminal Agent</title><link>https://www.neoteric.no/blog/worktree-isolation-crash-recovery-free-terminal-agent/</link><guid isPermaLink="true">https://www.neoteric.no/blog/worktree-isolation-crash-recovery-free-terminal-agent/</guid><description>Meta&apos;s Muse Code beta brings parallel sub-agents in isolated Git worktrees and a crash-safe event log, offering a free terminal agent alternative to Claude Code.</description><pubDate>Thu, 06 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>devtools</category><category>terminal</category><category>meta</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Retune Draft Models Without Restarting vLLM</title><link>https://www.neoteric.no/blog/retune-draft-models-without-restarting-vllm/</link><guid isPermaLink="true">https://www.neoteric.no/blog/retune-draft-models-without-restarting-vllm/</guid><description>vLLM v0.26.0 adds runtime draft-weight updates, so you can retune speculative decoding without restarting the engine and flushing the KV cache.</description><pubDate>Wed, 05 Aug 2026 09:00:00 GMT</pubDate><category>ai</category><category>open-source</category><category>devtools</category><category>vllm</category><category>inference</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Nova Deployments Are Going Away — Plan the Migration Now</title><link>https://www.neoteric.no/blog/nova-premier-omni-deprecation-migration/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nova-premier-omni-deprecation-migration/</guid><description>Amazon is deprecating Nova Premier, Omni, Reel and Canvas in a strategic pivot. Audit your Bedrock calls now, before a deadline lands.</description><pubDate>Tue, 04 Aug 2026 09:00:00 GMT</pubDate><category>industry-signal</category><category>cloud</category><category>ai</category><category>amazon-nova</category><category>bedrock</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>16,000 MT/s DDR5 MRDIMMs Just Raised the AI Memory Ceiling</title><link>https://www.neoteric.no/blog/mrdimm-16000-mts-ddr5/</link><guid isPermaLink="true">https://www.neoteric.no/blog/mrdimm-16000-mts-ddr5/</guid><description>Renesas&apos; Gen 3 MRDIMM pushes server DDR5 to 16,000 MT/s — 25% over Gen 2 on existing boards. Memory-bound decode gains a real ceiling.</description><pubDate>Tue, 04 Aug 2026 09:00:00 GMT</pubDate><category>memory</category><category>hardware</category><category>servers</category><category>homelab</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>GitHub Copilot SDK Powers Production Agents in MAF v1.0</title><link>https://www.neoteric.no/blog/github-copilot-sdk-agent-framework-v1/</link><guid isPermaLink="true">https://www.neoteric.no/blog/github-copilot-sdk-agent-framework-v1/</guid><description>Microsoft Agent Framework now backs agents with GitHub Copilot SDK, shipping shell execution, file ops, and MCP integration in stable C# and Python SDKs.</description><pubDate>Sun, 02 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>devtools</category><category>software-architecture</category><category>open-source</category><category>python</category><category>dotnet</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Agent Infrastructure Now Targets Memory Bottlenecks—PB-Scale Storage Changes Deployment</title><link>https://www.neoteric.no/blog/agent-memory-storage-pb-scale/</link><guid isPermaLink="true">https://www.neoteric.no/blog/agent-memory-storage-pb-scale/</guid><description>Huawei Cloud&apos;s new Agentic Infra adds PB-scale memory storage and unified scheduling. What this means for long-horizon agent tasks and where the real constraint sits.</description><pubDate>Sat, 01 Aug 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>infrastructure</category><category>cloud</category><category>business-agents</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Ruflo&apos;s MCP Bridge Exposes 233 Tools Without Auth—CVE-2026-59726</title><link>https://www.neoteric.no/blog/ruflo-mcp-bridge-cve-2026-59726/</link><guid isPermaLink="true">https://www.neoteric.no/blog/ruflo-mcp-bridge-cve-2026-59726/</guid><description>CVE-2026-59726 lets unauthenticated attackers execute shell commands, steal LLM keys, and poison agent memory via Ruflo&apos;s exposed /mcp endpoint. Upgrade to 3.16.3 and audit immediately.</description><pubDate>Fri, 31 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>ai</category><category>security</category><category>open-source</category><category>mcp</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Agents Now Call Physics Solvers Like Any Tool</title><link>https://www.neoteric.no/blog/agents-physics-solvers-cuda-toolkit/</link><guid isPermaLink="true">https://www.neoteric.no/blog/agents-physics-solvers-cuda-toolkit/</guid><description>NVIDIA embeds sparse linear algebra solvers into the Agent Toolkit—cuISS, cuDSS, cuEST—so autonomous design agents call physics simulation without leaving the flow. Free libraries, GPU-only execution.</description><pubDate>Mon, 27 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>ai</category><category>software-architecture</category><category>infrastructure</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>MCP Servers No Longer Need Session Affinity—Deploy Like Any Cloud App</title><link>https://www.neoteric.no/blog/mcp-stateless-scaling/</link><guid isPermaLink="true">https://www.neoteric.no/blog/mcp-stateless-scaling/</guid><description>MCP protocol drops session-based architecture July 28, making agents scale horizontally without sticky routing. New SDKs support both old and new versions, but sampling deprecation requires network auth changes.</description><pubDate>Sat, 25 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>software-architecture</category><category>cloud</category><category>devtools</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Server Memory Hit 70% Price Spike — Buy Before Q3 Peak</title><link>https://www.neoteric.no/blog/dram-price-spike-q3-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/dram-price-spike-q3-2026/</guid><description>SK Hynix chairman confirms 70% DRAM price hike. Conventional DRAM contract prices jumped 90–95% in Q1; Q3 will see another 58–63% increase. Homelab and server buyers need to act now.</description><pubDate>Wed, 22 Jul 2026 09:00:00 GMT</pubDate><category>memory</category><category>market-prices</category><category>hardware</category><category>dram</category><category>supply-chain</category><category>server</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>AMD EPYC Venice Launches July 22–24: 576 Cores, 12 TB/s Memory Bandwidth</title><link>https://www.neoteric.no/blog/amd-epyc-venice-launch-cores-memory-bandwidth/</link><guid isPermaLink="true">https://www.neoteric.no/blog/amd-epyc-venice-launch-cores-memory-bandwidth/</guid><description>AMD&apos;s Zen 6 EPYC Venice debuts next week with up to 576 cores, 12 TB/s memory bandwidth, and native AVX-1024 support. Here&apos;s what changes for AI infrastructure.</description><pubDate>Sat, 18 Jul 2026 09:00:00 GMT</pubDate><category>cpu</category><category>amd</category><category>epyc</category><category>server</category><category>ai</category><category>infrastructure</category><category>benchmarks</category><category>hardware</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Coding Agents Are Now a Direct RCE Vector</title><link>https://www.neoteric.no/blog/coding-agents-rce-mcp-supply-chain/</link><guid isPermaLink="true">https://www.neoteric.no/blog/coding-agents-rce-mcp-supply-chain/</guid><description>Straiker&apos;s STAR Labs found 36% of successful coding agent attacks achieve remote code execution on developer machines holding source code and cloud keys. MCP ecosystem unvetted.</description><pubDate>Wed, 15 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>security</category><category>ai</category><category>devtools</category><category>supply-chain</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Claude Code Governance Defaults Just Flipped</title><link>https://www.neoteric.no/blog/claude-code-auto-mode-default-bedrock-vertex/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-code-auto-mode-default-bedrock-vertex/</guid><description>Claude Code v2.1.207 removes the opt-in flag for auto mode on Bedrock, Vertex, and Azure—teams must now actively set disableAutoMode to maintain manual control before the update lands.</description><pubDate>Mon, 13 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>cloud</category><category>business-agents</category><category>software-architecture</category><category>devtools</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Homelab RAM and SSD Costs Are About to Jump 30–40%</title><link>https://www.neoteric.no/blog/q3-dram-nand-price-hike/</link><guid isPermaLink="true">https://www.neoteric.no/blog/q3-dram-nand-price-hike/</guid><description>ADATA&apos;s chairman warns of a 30% DRAM and 40% NAND jump this quarter. Here&apos;s why your build budget needs to shift now.</description><pubDate>Sun, 12 Jul 2026 09:00:00 GMT</pubDate><category>memory</category><category>market-prices</category><category>homelab</category><category>hardware</category><category>nand</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your AI Stack Choices Lock You Into Years of Vendor Control</title><link>https://www.neoteric.no/blog/sovereign-ai-five-dimensions/</link><guid isPermaLink="true">https://www.neoteric.no/blog/sovereign-ai-five-dimensions/</guid><description>Sovereign AI isn&apos;t data residency. It&apos;s territorial, operational, technological, legal—and financial control. The three case studies that changed CFO thinking in Q1 2026.</description><pubDate>Sun, 12 Jul 2026 09:00:00 GMT</pubDate><category>ai</category><category>software-architecture</category><category>cloud</category><category>infrastructure</category><category>business-agents</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>TypeScript 7 Cuts Builds 8–12×, But Vue and Svelte Wait Until 7.1</title><link>https://www.neoteric.no/blog/typescript-7-go-compiler-framework-gap/</link><guid isPermaLink="true">https://www.neoteric.no/blog/typescript-7-go-compiler-framework-gap/</guid><description>TypeScript 7.0 stable (July 8) ships a native Go compiler that slashes build times by 8–12× — but Vue, Svelte, and Astro lose template type-checking until the programmatic API stabilizes in 7.1.</description><pubDate>Sat, 11 Jul 2026 09:00:00 GMT</pubDate><category>devtools</category><category>software-architecture</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Stacking RL on Already-RL Models Breaks the Ceiling</title><link>https://www.neoteric.no/blog/rl-stacking-breaks-ceiling/</link><guid isPermaLink="true">https://www.neoteric.no/blog/rl-stacking-breaks-ceiling/</guid><description>Cognition&apos;s SWE-1.7 adds 12.2 points atop an already-trained base, challenging the field&apos;s assumption that RL post-training hits a hard wall.</description><pubDate>Fri, 10 Jul 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>benchmarks</category><category>software-architecture</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>295B Parameters, 21B Active—Hy3 Proves Size Isn&apos;t Agent Performance</title><link>https://www.neoteric.no/blog/hy3-smaller-model-agent-tasks/</link><guid isPermaLink="true">https://www.neoteric.no/blog/hy3-smaller-model-agent-tasks/</guid><description>Tencent&apos;s Hy3 benchmarks between models 2–5x larger by parameter count. 256K context, Apache 2.0, 1 yuan per million input tokens. The catch: vendor benchmarks need independent testing.</description><pubDate>Thu, 09 Jul 2026 09:00:00 GMT</pubDate><category>ai</category><category>agentic-workflows</category><category>open-source</category><category>benchmarks</category><category>local-ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Task-Specific Logic Compiles to 23MB Files Running Offline on 600M Models</title><link>https://www.neoteric.no/blog/paw-offline-adapters-32b-equivalent/</link><guid isPermaLink="true">https://www.neoteric.no/blog/paw-offline-adapters-32b-equivalent/</guid><description>University of Waterloo&apos;s PAW system compiles fuzzy-function specs into 23MB LoRA adapters that match 32B model accuracy on a 600M interpreter—no API calls, no cloud dependency, 30 tok/s on MacBook M3.</description><pubDate>Sat, 04 Jul 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>agentic-workflows</category><category>open-source</category><category>software-architecture</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Agents Must Now Reuse Your Design System</title><link>https://www.neoteric.no/blog/a2ui-v0-9-design-system-agents/</link><guid isPermaLink="true">https://www.neoteric.no/blog/a2ui-v0-9-design-system-agents/</guid><description>A2UI v0.9 flips the script: agents declare UI intent against your existing design catalog, not invent components. What changes for agentic architectures.</description><pubDate>Fri, 03 Jul 2026 09:00:00 GMT</pubDate><category>agentic-workflows</category><category>ai</category><category>software-architecture</category><category>devtools</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>RTX 50 Prices Jump 19% and the 5070 Ti Is Gone — Buy Now or Wait</title><link>https://www.neoteric.no/blog/nvidia-rtx-50-shortage-price-hike/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-rtx-50-shortage-price-hike/</guid><description>NVIDIA cut RTX 50 supply by 20% and sidelined the 5070 Ti. Street prices are up 19%. Here&apos;s how it changes your local AI build.</description><pubDate>Thu, 25 Jun 2026 09:00:00 GMT</pubDate><category>nvidia</category><category>gpu</category><category>market-prices</category><category>local-ai</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>GLM-5.2 Shifts Agentic Coding Economics to 1/6th Frontier Cost</title><link>https://www.neoteric.no/blog/glm-5-2-agentic-coding-economics/</link><guid isPermaLink="true">https://www.neoteric.no/blog/glm-5-2-agentic-coding-economics/</guid><description>Z.AI&apos;s open-weights GLM-5.2 hits 81% on Terminal-Bench, undercutting GPT-5.5 for a fraction of the cost. Here&apos;s the benchmark breakdown.</description><pubDate>Sat, 20 Jun 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>benchmarks</category><category>glm-5-2</category><category>agentic-workflows</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>DiffusionGemma Puts 1000+ tok/s on an RTX 5090</title><link>https://www.neoteric.no/blog/diffusiongemma-rtx-5090-vram/</link><guid isPermaLink="true">https://www.neoteric.no/blog/diffusiongemma-rtx-5090-vram/</guid><description>Google&apos;s open 26B diffusion model hits 700+ tok/s on consumer GPUs with day-zero vLLM support. Here&apos;s what the bidirectional architecture changes for local inference.</description><pubDate>Sat, 13 Jun 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>benchmarks</category><category>diffusiongemma</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>The End of On-Demand Frontier AI</title><link>https://www.neoteric.no/blog/the-end-of-on-demand-frontier-ai/</link><guid isPermaLink="true">https://www.neoteric.no/blog/the-end-of-on-demand-frontier-ai/</guid><description>Google&apos;s $920M/month SpaceX contract shifts AI infrastructure from on-demand cloud to locked-in capacity. Here&apos;s what changes for enterprise architecture.</description><pubDate>Fri, 12 Jun 2026 09:29:00 GMT</pubDate><category>industry-signal</category><category>ai-infrastructure</category><category>google</category><category>spacex</category><category>cloud-compute</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>The Grid, Not the GPU, Is Your Next Inference Bottleneck</title><link>https://www.neoteric.no/blog/the-grid-not-the-gpu-is-your-next-inference-bottleneck/</link><guid isPermaLink="true">https://www.neoteric.no/blog/the-grid-not-the-gpu-is-your-next-inference-bottleneck/</guid><description>Seattle&apos;s data center ban and 3-year grid queues mean local inference and modular hardware are no longer optional.</description><pubDate>Fri, 12 Jun 2026 09:28:00 GMT</pubDate><category>ai-infrastructure</category><category>local-ai</category><category>modular-datacenters</category><category>energy</category><category>grid-constraints</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>NVIDIA Halves Vera Rubin Memory — LPDDR5X Shortage Hits First</title><link>https://www.neoteric.no/blog/nvidia-vera-rubin-memory-cut-supply/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-vera-rubin-memory-cut-supply/</guid><description>NVIDIA&apos;s next-gen Vera Rubin drops standard CPU memory from 54TB to 28TB per rack. Here&apos;s what the LPDDR5X shortage means for AI infrastructure.</description><pubDate>Fri, 12 Jun 2026 09:00:00 GMT</pubDate><category>nvidia</category><category>memory</category><category>industry-signal</category><category>supply-chain</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>CI/CD Workflows Just Leaked Secrets to Untrusted Prompts</title><link>https://www.neoteric.no/blog/claude-code-github-action-secret-leak/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-code-github-action-secret-leak/</guid><description>Microsoft found Claude Code&apos;s GitHub Action exposes runner secrets via the Read tool. Here&apos;s the exact attack path and how to lock it down.</description><pubDate>Fri, 12 Jun 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>agentic-workflows</category><category>security</category><category>ci-cd</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>OpenAI&apos;s $85B Burn Rate Changes the AI Cost Equation</title><link>https://www.neoteric.no/blog/openai-s-85b-burn-rate-changes-the-ai-cost-equation/</link><guid isPermaLink="true">https://www.neoteric.no/blog/openai-s-85b-burn-rate-changes-the-ai-cost-equation/</guid><description>The SEC filing exposes a nine-to-one burn-to-revenue ratio. Here is why local inference is shifting from a privacy play to a hard P&amp;L requirement.</description><pubDate>Thu, 11 Jun 2026 15:26:00 GMT</pubDate><category>ai</category><category>openai</category><category>ipo</category><category>ai-infrastructure</category><category>local-ai</category><category>enterprise-ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your API Costs Are About to Correct</title><link>https://www.neoteric.no/blog/your-api-costs-are-about-to-correct/</link><guid isPermaLink="true">https://www.neoteric.no/blog/your-api-costs-are-about-to-correct/</guid><description>OpenAI&apos;s IPO filing ends the era of subsidized frontier AI. Here&apos;s what the pricing shift means for local inference, enterprise contracts, and your agentic stack.</description><pubDate>Thu, 11 Jun 2026 11:50:00 GMT</pubDate><category>ai</category><category>openai</category><category>local-ai</category><category>enterprise-ai</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Low-Code AI Automations: How Gartner Sees the Future of Development</title><link>https://www.neoteric.no/blog/low-code-ai-automations-how-gartner-sees-the-future-of-development/</link><guid isPermaLink="true">https://www.neoteric.no/blog/low-code-ai-automations-how-gartner-sees-the-future-of-development/</guid><description>By 2026, 70-75% of enterprise apps will use low-code platforms. AI-powered low-code will drive 80% of app development by 2029, mainly from non-IT users. Yet 40% of agentic AI projects face cancellation due to governance gaps. OutSystems, Mendix, and Appian lead the market. Success requires training, governance, and embedded security.</description><pubDate>Fri, 05 Jun 2026 09:00:00 GMT</pubDate><category>low-code</category><category>ai</category><category>automation</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>NVIDIA RTX Spark: What 128 GB of Unified Memory Means for Local AI</title><link>https://www.neoteric.no/blog/nvidia-rtx-spark-local-ai/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-rtx-spark-local-ai/</guid><description>NVIDIA&apos;s first consumer SoC brings 128 GB of unified memory and native CUDA to Windows laptops — here&apos;s what it means for running LLMs locally, and the catches.</description><pubDate>Wed, 03 Jun 2026 09:09:00 GMT</pubDate><category>nvidia</category><category>local-ai</category><category>rtx-spark</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Apple Silicon Just Cracked the 150B Wall at 56 tok/s</title><link>https://www.neoteric.no/blog/rapid-mlx-deepseek-v4-flash-158b/</link><guid isPermaLink="true">https://www.neoteric.no/blog/rapid-mlx-deepseek-v4-flash-158b/</guid><description>Rapid-MLX&apos;s new engine optimizations push DeepSeek V4 Flash 158B-A13B to 56 tok/s on a Mac. Here are the exact numbers and the tradeoffs.</description><pubDate>Tue, 02 Jun 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>mlx</category><category>benchmarks</category><category>moe</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Choosing a Messaging System: A 2026 Architect&apos;s Guide</title><link>https://www.neoteric.no/blog/mqtt-azure-service-bus-event-grid-event-hub-amazon-sqs-messaging-system-comparison/</link><guid isPermaLink="true">https://www.neoteric.no/blog/mqtt-azure-service-bus-event-grid-event-hub-amazon-sqs-messaging-system-comparison/</guid><description>Message queues, brokers, and event streams — and how to pick the right one.</description><pubDate>Tue, 02 Jun 2026 00:07:00 GMT</pubDate><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>NVIDIA&apos;s First Server CPU Clears the Performance Floor</title><link>https://www.neoteric.no/blog/nvidia-vera-cpu-performance-floor/</link><guid isPermaLink="true">https://www.neoteric.no/blog/nvidia-vera-cpu-performance-floor/</guid><description>NVIDIA&apos;s Vera CPU benchmarks show it competes directly with EPYC and Xeon, but software maturity and platform lock-in dictate adoption.</description><pubDate>Sat, 30 May 2026 09:00:00 GMT</pubDate><category>hardware</category><category>cpu</category><category>nvidia</category><category>ai</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>On-Device Agents Just Gained a 6GB MoE That Actually Works</title><link>https://www.neoteric.no/blog/on-device-agents-6gb-moe/</link><guid isPermaLink="true">https://www.neoteric.no/blog/on-device-agents-6gb-moe/</guid><description>Liquid AI&apos;s LFM2.5-8B-A1B hits under 6GB RAM with 1.5B active parameters, native tool calling, and verified throughput across edge hardware.</description><pubDate>Sat, 30 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>on-device-ai</category><category>agentic-workflows</category><category>liquid-ai</category><category>mlx</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your 12GB MTP Throughput Just Jumped 23%</title><link>https://www.neoteric.no/blog/12gb-mtp-throughput-jumps-23-percent/</link><guid isPermaLink="true">https://www.neoteric.no/blog/12gb-mtp-throughput-jumps-23-percent/</guid><description>A community llama.cpp fork squeezes 110 tok/s out of Qwen3.6-35B-A3B MTP on a 12GB card. Here are the exact flags and the VRAM trick to make it fit.</description><pubDate>Fri, 29 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>llama-cpp</category><category>benchmarks</category><category>mtp</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Local Throughput Just Doubled (With No Accuracy Tax)</title><link>https://www.neoteric.no/blog/llama-cpp-mtp-speculative-decoding/</link><guid isPermaLink="true">https://www.neoteric.no/blog/llama-cpp-mtp-speculative-decoding/</guid><description>Multi-token prediction is merged into llama.cpp, delivering 1.4–2.2x throughput on Qwen3.6 with zero accuracy loss.</description><pubDate>Tue, 26 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>llama-cpp</category><category>benchmarks</category><category>homelab</category><category>speculative-decoding</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Workstations Get 192GB RAM and 160GB VRAM for AI in Q3</title><link>https://www.neoteric.no/blog/workstation-ram-vram-ai-q3/</link><guid isPermaLink="true">https://www.neoteric.no/blog/workstation-ram-vram-ai-q3/</guid><description>AMD&apos;s next-gen Ryzen AI Max chips bring desktop AI consolidation, massive unified memory, and high NPU TOPS to compact workstations starting this fall.</description><pubDate>Tue, 26 May 2026 09:00:00 GMT</pubDate><category>cpu</category><category>memory</category><category>chipset</category><category>workstation</category><category>ai</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Remote MCP Timeouts Just Stopped Hitting 60 Seconds</title><link>https://www.neoteric.no/blog/remote-mcp-timeouts-fix/</link><guid isPermaLink="true">https://www.neoteric.no/blog/remote-mcp-timeouts-fix/</guid><description>Claude Code v2.1.149 breaks the hard 60s cap on remote tool calls and adds session listing for scripting.</description><pubDate>Mon, 25 May 2026 09:00:00 GMT</pubDate><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>MTP in llama.cpp Adds 24–33% Throughput on Qwen3</title><link>https://www.neoteric.no/blog/llamacpp-mtp-qwen3-throughput-gain/</link><guid isPermaLink="true">https://www.neoteric.no/blog/llamacpp-mtp-qwen3-throughput-gain/</guid><description>llama.cpp merged Multi-Token Prediction for Qwen3. Community benchmarks show 38→47 tok/s on RTX 3090 and 63→84 tok/s on RTX 5090 — no new hardware needed.</description><pubDate>Sun, 24 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>benchmarks</category><category>llama-cpp</category><category>agentic-workflows</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>AMD GPUs Are Getting 10% Pricier — Here&apos;s the Full Picture</title><link>https://www.neoteric.no/blog/amd-gpu-price-hike-dram-shortage-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/amd-gpu-price-hike-dram-shortage-2026/</guid><description>AMD has notified its supply chain of a ~10% GPU price increase driven by a DRAM shortage. DRAM contract prices rose 90–95% QoQ in Q1 2026. Here&apos;s what to buy, hold, or skip.</description><pubDate>Sat, 23 May 2026 09:00:00 GMT</pubDate><category>hardware</category><category>gpu</category><category>memory</category><category>market-prices</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>DFlash Decode Collapses at Long Output. MTP Doesn&apos;t.</title><link>https://www.neoteric.no/blog/llama-cpp-mtp-spec-draft-p-min-long-context/</link><guid isPermaLink="true">https://www.neoteric.no/blog/llama-cpp-mtp-spec-draft-p-min-long-context/</guid><description>One llama.cpp flag—--spec-draft-p-min 0.75—turns MTP from a dud into a decode speed that holds flat across output lengths where DFlash falls apart.</description><pubDate>Sat, 23 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>benchmarks</category><category>homelab</category><category>llama-cpp</category><category>speculative-decoding</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Code&apos;s /simplify Stopped Fixing Code Yesterday</title><link>https://www.neoteric.no/blog/claude-code-s-simplify-stopped-fixing-code-yesterday/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-code-s-simplify-stopped-fixing-code-yesterday/</guid><description>Claude Code 2.1.147 renamed /simplify to /code-review and dropped the auto-fix behavior. The new command reports bugs at chosen effort levels but no longer changes code.</description><pubDate>Fri, 22 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>agentic-workflows</category><category>tooling</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>DDR5 Tripled Since October — Buy Nothing Yet</title><link>https://www.neoteric.no/blog/dram-nand-price-crunch-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/dram-nand-price-crunch-2026/</guid><description>DRAM and NAND prices have surged 90–130% since late 2025. A Samsung factory walkout and HBM reallocation mean the crunch won&apos;t clear before late 2027 at earliest.</description><pubDate>Fri, 22 May 2026 09:00:00 GMT</pubDate><category>hardware</category><category>memory</category><category>market-prices</category><category>homelab</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>One llama.cpp Flag Turns MTP From Dead Weight to 68% Faster</title><link>https://www.neoteric.no/blog/llamacpp-spec-draft-p-min-mtp-qwen3/</link><guid isPermaLink="true">https://www.neoteric.no/blog/llamacpp-spec-draft-p-min-mtp-qwen3/</guid><description>The --spec-draft-p-min filter in llama.cpp PR #22397 rescues MTP for Qwen3.6-27B: 48.9 tok/s vs 29 tok/s at 2000 tokens on a 24GB card.</description><pubDate>Fri, 22 May 2026 09:00:00 GMT</pubDate><category>local-ai</category><category>llama-cpp</category><category>benchmarks</category><category>qwen</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Private MCP Server Is Now Claude-Reachable</title><link>https://www.neoteric.no/blog/your-private-mcp-server-is-now-claude-reachable/</link><guid isPermaLink="true">https://www.neoteric.no/blog/your-private-mcp-server-is-now-claude-reachable/</guid><description>Anthropic shipped MCP tunnels on May 19. Claude agents can call internal databases, ticketing systems, and on-prem APIs through one outbound connection — no inbound firewall rules required.</description><pubDate>Thu, 21 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>mcp</category><category>agentic-workflows</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your vLLM Thinking Budget Was Doing Nothing With MTP On</title><link>https://www.neoteric.no/blog/your-vllm-thinking-budget-was-doing-nothing-with-mtp-on/</link><guid isPermaLink="true">https://www.neoteric.no/blog/your-vllm-thinking-budget-was-doing-nothing-with-mtp-on/</guid><description>vLLM 0.21.0 shipped Friday with a quiet fix: thinking_token_budget was being silently ignored when MTP speculative decoding was enabled. If you serve reasoning models with spec decode, you have been paying for it.</description><pubDate>Mon, 18 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>vllm</category><category>speculative-decoding</category><category>benchmarks</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Code v2.1.100+ Burns ~20K Phantom Tokens Per Request</title><link>https://www.neoteric.no/blog/claude-code-v2-1-100-burns-20k-phantom-tokens-per-request/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-code-v2-1-100-burns-20k-phantom-tokens-per-request/</guid><description>A server-side bug in Claude Code v2.1.100+ inflates every request by roughly 20K cache_creation tokens — about 40% overhead. Pin v2.1.98 until fixed.</description><pubDate>Sun, 17 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>agentic-workflows</category><category>benchmarks</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Your Local Qwen3.6 Throughput Probably Just Halved (and How to Fix It)</title><link>https://www.neoteric.no/blog/llama-cpp-mtp-flag-rename/</link><guid isPermaLink="true">https://www.neoteric.no/blog/llama-cpp-mtp-flag-rename/</guid><description>llama.cpp renamed the MTP flag on May 13. The old --spec-type mtp is silently ignored. If your tok/s dropped from 140 to 70 you are likely running without speculative decoding.</description><pubDate>Sat, 16 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>llama-cpp</category><category>qwen</category><category>local-ai</category><category>speculative-decoding</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>MCP Server Roundup: Which Are Actually Worth Adding to Your Setup in May 2026</title><link>https://www.neoteric.no/blog/mcp-server-roundup-may-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/mcp-server-roundup-may-2026/</guid><description>Eighteen months after Anthropic released MCP, the ecosystem is wide enough that picking the wrong servers slows your agent down. Here is the practical short list — what to install, what to skip, and the trap most people fall into.</description><pubDate>Sat, 16 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>mcp</category><category>claude</category><category>tooling</category><category>agentic-workflows</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Speculative Decoding Explained: Why Your Local Model Got 2× Faster in 2026</title><link>https://www.neoteric.no/blog/speculative-decoding-why-local-ai-got-fast/</link><guid isPermaLink="true">https://www.neoteric.no/blog/speculative-decoding-why-local-ai-got-fast/</guid><description>The same Qwen3.6-27B that ran at 70 tokens/sec on a 4090 in January was running at 140 tokens/sec by April. Nothing changed about the model. Speculative decoding moved from research curiosity to default. Here is what it actually does.</description><pubDate>Sat, 16 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>llama-cpp</category><category>performance</category><category>speculative-decoding</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Third-Party Claude Agents Lose the Subscription Subsidy June 15</title><link>https://www.neoteric.no/blog/third-party-claude-agents-lose-the-subscription-subsidy-june-15/</link><guid isPermaLink="true">https://www.neoteric.no/blog/third-party-claude-agents-lose-the-subscription-subsidy-june-15/</guid><description>Anthropic is splitting Claude billing on June 15 — Agent SDK and ACP usage moves to a capped credit pool ($20/$100/$200) at full API rates.</description><pubDate>Sat, 16 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>agentic-workflows</category><category>industry-signal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>The Local AI Inflection Point: May 2026</title><link>https://www.neoteric.no/blog/local-ai-inflection-point-may-2026/</link><guid isPermaLink="true">https://www.neoteric.no/blog/local-ai-inflection-point-may-2026/</guid><description>Three model releases in three weeks moved local AI from &apos;good enough for hobbies&apos; to &apos;good enough for production&apos;. Here&apos;s what changed and why it matters.</description><pubDate>Fri, 15 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>local-ai</category><category>qwen</category><category>gemma</category><category>self-hosted</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Running Qwen3.6-27B Locally: Hardware, Quantization, and What Actually Works</title><link>https://www.neoteric.no/blog/running-qwen-3-6-27b-locally/</link><guid isPermaLink="true">https://www.neoteric.no/blog/running-qwen-3-6-27b-locally/</guid><description>A practical guide to running Qwen3.6-27B on consumer hardware in 2026 — memory requirements per quant level, recommended runners, and the MTP trick that doubles your tokens per second.</description><pubDate>Mon, 11 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>qwen</category><category>local-ai</category><category>llama-cpp</category><category>homelab</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>A 27B Model on a Single GPU Is 10 Points Off Claude Opus 4.7</title><link>https://www.neoteric.no/blog/qwen-3-6-27b-vs-claude-opus-4-7-benchmarks/</link><guid isPermaLink="true">https://www.neoteric.no/blog/qwen-3-6-27b-vs-claude-opus-4-7-benchmarks/</guid><description>Qwen3.6-27B running locally now scores within 10 points of frontier closed models on SWE-bench Verified. The benchmark table, lined up side by side.</description><pubDate>Fri, 08 May 2026 09:00:00 GMT</pubDate><category>ai</category><category>qwen</category><category>claude</category><category>local-ai</category><category>benchmarks</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Opus 4.7: 87.6% on SWE-bench and 1M Context at Standard Pricing</title><link>https://www.neoteric.no/blog/claude-opus-4-7-coding-leap/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-opus-4-7-coding-leap/</guid><description>Anthropic shipped Opus 4.7 on April 16, 2026, with a seven-point SWE-bench jump, the 1M context window now generally available with no premium, and a new task budget primitive for agent loops.</description><pubDate>Fri, 17 Apr 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>anthropic</category><category>coding</category><category>agents</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Gemma 4: Google&apos;s Open Model Family Goes Multimodal</title><link>https://www.neoteric.no/blog/google-gemma-4-open-models/</link><guid isPermaLink="true">https://www.neoteric.no/blog/google-gemma-4-open-models/</guid><description>Google released Gemma 4 on April 2, 2026 — four variants from 2B to 31B, with 256K context, native vision and audio, and Apache 2.0 licensing. Here&apos;s what it&apos;s for, where it fits, and how to run it.</description><pubDate>Sun, 05 Apr 2026 09:00:00 GMT</pubDate><category>ai</category><category>google</category><category>gemma</category><category>local-ai</category><category>multimodal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>What&apos;s New in Optimizely CMS 13: The Big Picture</title><link>https://www.neoteric.no/blog/optimizely-cms-13-whats-new/</link><guid isPermaLink="true">https://www.neoteric.no/blog/optimizely-cms-13-whats-new/</guid><description>Optimizely CMS 13 went GA on April 1, 2026. Visual Builder is now the default editor, Content Manager replaces tree-first navigation, Optimizely Graph and Opti ID are mandatory, and the platform jumps to .NET 10. Here&apos;s what actually changed, where it&apos;s worth caring, and what the upgrade is going to cost you.</description><pubDate>Wed, 01 Apr 2026 09:00:00 GMT</pubDate><category>optimizely</category><category>cms</category><category>dotnet</category><category>headless</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Opus 4.6: A Million-Token Context and a New Agent Team Model</title><link>https://www.neoteric.no/blog/claude-opus-4-6-million-token-context-and-agent-teams/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-opus-4-6-million-token-context-and-agent-teams/</guid><description>Anthropic released Opus 4.6 on February 5, 2026, with a 1M token context beta, agent teams, adaptive thinking, and developer effort controls — all at the same price as 4.5.</description><pubDate>Fri, 06 Feb 2026 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>anthropic</category><category>agents</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Introducing Azure DevOps Workflow: Manage Work Items Without Leaving VS Code</title><link>https://www.neoteric.no/blog/azure-devops-workflow-vscode-extension/</link><guid isPermaLink="true">https://www.neoteric.no/blog/azure-devops-workflow-vscode-extension/</guid><description>A VS Code extension that brings Azure DevOps sprint boards, work item management, and AI-powered assistance directly into your editor.</description><pubDate>Tue, 25 Nov 2025 09:00:00 GMT</pubDate><category>vscode</category><category>azure-devops</category><category>productivity</category><category>open-source</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Claude Opus 4.5: Anthropic&apos;s New Flagship Model Sets the Bar for AI Coding</title><link>https://www.neoteric.no/blog/claude-opus-4-5-anthropics-new-flagship/</link><guid isPermaLink="true">https://www.neoteric.no/blog/claude-opus-4-5-anthropics-new-flagship/</guid><description>Anthropic&apos;s latest model achieves state-of-the-art results in agentic coding and brings meaningful improvements across reasoning, mathematics, and everyday tasks.</description><pubDate>Tue, 25 Nov 2025 09:00:00 GMT</pubDate><category>ai</category><category>claude</category><category>anthropic</category><category>coding</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item><item><title>Google Gemini 3 Pro: The New Leader in Multimodal AI</title><link>https://www.neoteric.no/blog/google-gemini-3-pro-multimodal-reasoning/</link><guid isPermaLink="true">https://www.neoteric.no/blog/google-gemini-3-pro-multimodal-reasoning/</guid><description>Google&apos;s Gemini 3 Pro brings generative interfaces, 1M token context, and state-of-the-art multimodal reasoning to developers and consumers alike.</description><pubDate>Tue, 25 Nov 2025 09:00:00 GMT</pubDate><category>ai</category><category>google</category><category>gemini</category><category>multimodal</category><author>post@neoteric.no (Hans Christian Thjømøe)</author></item></channel></rss>