Related posts
Sebastian Raschka
x.com
Big overhaul on DeepSeek V4.1 using an encoder-decoder setup. Tbh they should have called it DeepSeek V5! Super cool and refreshing, though!
overhaul encoder decoder
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
Writing
x.com Blog GitHub
Top people
Sebastian Raschka
Matt Mullenweg
John Resig
Salvatore Sanfilippo
Raph Koster
Lilian Weng
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
26 August
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers; - a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B; - a DeepSeek V4-style mHC residual path with four parallel streams; - plus a native vision encoder (not shown). * "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA +… Quoting @Zai_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now… LLMs Related
7 August
Co-founder of WordPress and founder of Automattic; blogs at ma.tt.
Toni on Verge Toni Schneider, Automattic’s founding CEO, board member, and now CEO of Bluesky, has a great conversation with Nilay Patel on The Verge’s Decoder podcast (YouTube, Pocket Casts). Automattic invested in Bluesky back in 2…
Related
4 June
JavaScript programmer; creator of the jQuery library, author of Secrets of the JavaScript Ninja, and long-time engineer at Khan Academy.
My pie-in-the-sky requests for @CloudflareDev to overhaul corporate web dev: * Fully hosted private Git + CI + "Pull Requests" * Playwright end-to-end test running and result collection * Storybook hosting + shapshot comparisons * Private NPM package registry hosting Related
15 February
Programmer; wrote Redis and hping, and blogs about C, systems programming and working alone.
29 October 2025
Game designer; wrote A Theory of Fun for Game Design and worked on Ultima Online and Star Wars Galaxies.
Stars Reach visual upgrades For months now, we have been working on a big overhaul to the visuals in Stars Reach. This has involved redoing most of the shaders in the entire game, and going back and retouching every asset. We started by redoing al…
Related
9 June 2022
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024, latterly leading its safety systems team.
Generalized Visual Language Models Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visua…
Related
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it