Related posts
Salvatore Sanfilippo
GitHub
llama.cpp-deepseek-v4-flash — Experimental implementation of DeepSeek v4 flaash in llama.cpp
llama flash deepseek
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
GitHub x.com Newsletter Korrents Recommends Bluesky Site
Top people
Sebastian Raschka
Thomas Wolf
Dylan Patel
Nathan Lambert
Salvatore Sanfilippo
Guillermo Rauch
Andrew Chen
Ismail Ghallou
Kun Chen
Aris Ripandi
Keyu Jin
Ethan Mollick
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
20 September
Programmer; wrote Redis and hping, and blogs about C, systems programming and working alone.
19 September
CEO of Vercel; creator of Next.js and Socket.IO. Writes at rauchg.com.
Looks like today may be a record day for token volume % of open models on Vercel AI Gateway: 🟦 Open 78.4% 🟨 Closed 21.6% While spend 💲 usually tells a different story, #3 and #4 today are Moonshot AI & DeepSeek. Adding Z.ai, their combined spend surpasses OpenAI (#2). (Do note that's the spend for inference of the model across providers (mostly in the US), not revenue going directly to the open weight labs.) OpenAI Related
13 September
General partner at Andreessen Horowitz and author of The Cold Start Problem. Previously led rider growth teams at Uber; has blogged on growth, network effects and marketplaces since 2007.
current homelab setup for local AI experimentation: - hermes box hosted on a Framework Desktop Mainboard AI Max+ 395 - 5090 eGPU running Qwen 3.8 27B for fast tok/s LLM use - sometimes 150+ tok/s - 2x DGX Spark: running Deepseek v4 Flash 0731 - better but slower model - pi 5 for monitoring - Mac mini as a dev box - use Herdr and ohmypi/codex/claude depending on the use case - housed in a 10" DeskPi mini rack (mostly) Hermes is defaulted to local AI but with a homegrown routing plugin hitting a small low TTFT model (Arch-Router) to decide whether to go local or upgrade to cloud/frontier. Tryin… LLMs Anthropic Related
Moroccan senior frontend developer writing at smakosh.com; builds side projects and client work under Smakosh LLC.
Even better guys, get a @LLMDevPass and let Astra plan and DeepSeek V4.1 Flash and GLM-5.3 Flash execute. You could do a lot through custom agents by having one senior agent handing over and delegating tasks to other sub agents, all of this is built-in @Empryo_ai Quoting @addyosmani Get more out of your Fable usage by using Opus for subagents. Great tip! Related
10 September
Former L8 engineer at Meta, Microsoft and Atlassian, now writing and building solo on agentic engineering. Writes Kun's Field Notes and posts a lot about AI coding agents on X.
omfg Qwen, GLM, Kimi, DeepSeek looks like they are all just claude Quoting @AnthropicAI We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons… Anthropic Related
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
Big overhaul on DeepSeek V4.1 using an encoder-decoder setup. Tbh they should have called it DeepSeek V5! Super cool and refreshing, though! Quoting @deepseek_ai 🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6 Related
26 August
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers; - a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B; - a DeepSeek V4-style mHC residual path with four parallel streams; - plus a native vision encoder (not shown). * "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA +… Quoting @Zai_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now… LLMs Related
16 August
Indonesian software engineer, educator and engineering manager at Zero One Group; open-source enthusiast writing at ripandis.com.
16 May
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
13 March
Founder and chief analyst of SemiAnalysis; reports on semiconductor supply chains, AI datacentre economics and what the chip export controls actually do.
Their words
So when you look at inference at let's say 100 tokens a second for deepseek and kimk 2.5 hopper versus blackwell the performance difference is on the order of 20x
Show the whole quote
youtube.com
30 December 2025
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
3 December 2025
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
13 August 2025
Economist specialising in China's economy, international macroeconomics and global trade, and author of The New China Playbook.
China America
Their words
And remember that DeepSeek happened in times of crisis, urgency, not in times of comfort. A lot of these technological breakthroughs and leapfrogging happens in times of crisis. This is called crisis innovation. And you got to thank the US for that. When the Chinese were comfortably importing chips from the US, the whole industry stalled for 20 years.
Show the whole quote
youtube.com
19 July 2025
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
23 June 2025
Professor at Wharton studying how people actually use AI, and author of Co-Intelligence. Writes One Useful Thing, where the claims are dated and testable because the thing they describe keeps changing under them.
Their words
r1, a Chinese model, is very capable and free to use, but is missing a few features from the other companies and it is not clear that they will keep up in the long term.
Show the whole quote
oneusefulthing.org
31 March 2025
Co-founder and chief science officer of Hugging Face; writes about open models and where he thinks the field is wrong.
Wow - open-source LLMs for code are so back Just one week after the latest DeepSeek release, here is the All-Hands team dropping a 32B model with matching performance on software engineering tasks benchmarks (SWE-Bench). 20 times smaller than the DeepSeek model -- LLMs benchmarks Related
12 March 2025
Co-founder and chief science officer of Hugging Face; writes about open models and where he thinks the field is wrong.
We've kept pushing our Open-R1 project, an open initiative to replicate and extend the techniques behind DeepSeek-R1 And even we were mind-blown by the results we got with this latest model we're releasing: ⚡️OlympicCoder [1/3] Related
27 February 2025
Co-founder and chief science officer of Hugging Face; writes about open models and where he thinks the field is wrong.
A few words on DeepSeek new releases. Links are: - github.com/deepseek-ai/... - github.com/deepseek-ai/... - github.com/deepseek-ai/... and the Ultra-Scale Playbook at huggingface.co
Related
3 February 2025
Machine-learning researcher on open language models; writes the Interconnects newsletter and the RLHF Book, after leading post-training at Ai2.
From one piece
DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459
2 beliefs, in the piece's order there
OpenAI startups
Their words
I think that they're trying to shift the narrative. They're trying to protect themselves. We saw this years ago when ByteDance was actually banned from some OpenAI APIs for training on outputs. There's other AI startups that most people, if you're in the AI culture, were like they just told us they trained on OpenAI outputs and they never got banned.
Show the whole quote
youtube.com
Their words
The more progress that AI makes or the higher the derivative of AI progress is, especially because NVIDIA's in the best place, the higher the derivative is, the sooner the market's going to be bigger and expanding and NVIDIA's the only one that does everything reliably right now.
Show the whole quote
youtube.com
Founder and chief analyst of SemiAnalysis; reports on semiconductor supply chains, AI datacentre economics and what the chip export controls actually do.
Their words
But the funniest thing I think that comes out of this is Jevons paradox is true. AWS pricing for H100s has gone up over the last couple of weeks, since a little bit after Christmas, since V3 was launched, AWS H100 pricing has gone up.
Show the whole quote
youtube.com
2 February 2025
Independent AI policy researcher; led policy research at OpenAI from 2018 to 2024, latterly as senior adviser for AGI readiness.
30 January 2025
Creator of Flask and Jinja. Writes about software at lucumr.pocoo.org.
Their words
First and foremost, I use this to talk to local models hosted by Ollama, but secondarily I also use it to interface with other remote services like OpenAI, Anthropic and DeepSeek.
Show the whole quote
lucumr.pocoo.org
26 January 2025
Design engineer and illustrator; makes visual essays on programming, anthropology and what language models do to the way people write.
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it