Related posts
Sebastian Raschka
x.com
Big overhaul on DeepSeek V4.1 using an encoder-decoder setup. Tbh they should have called it DeepSeek V5! Super cool and refreshing, though!
overhaul encoder refreshing
The the words it uses that this site has seen
least often elsewhere. Posts are matched on those words alone —
nothing here is a summary of this one.
Sources
All
x.com Bluesky Recommends Blog
Top people
Hamel Husain
Gergely Orosz
Sebastian Raschka
Saloni Dattani
John Resig
Jason Furman
Raph Koster
Zvi Mowshowitz
Paul Millerd
Vicki Boykis
Lilian Weng
Showing
Profile →
Show everything
Hiding
Show them again
Show them again
Further back ↓
Hiding
Show them again
30 August
Machine learning engineer and independent AI consultant. He writes about LLM evaluation, tooling and applied ML at hamel.dev, and previously worked on machine learning at GitHub.
Teaching is hard but worth it when you see this 🥰 @sh_reya and I are refreshing the material once again to incorporate the latest eval techniques in our next cohort Related
26 August
Writes The Pragmatic Engineer, the software-engineering newsletter, and wrote The Software Engineer's Guidebook. Formerly an engineering manager at Uber, and at Microsoft/Skype and Skyscanner before that.
Few people care more about performant code than @cmuratori.bsky.social. Casey is a down-to-earth programmer, game dev, performance nerd. A refreshing episode about why the basics (still) matter, and why learning to read Assembly remains useful for anyone who cares about performance: (cont'd) Related
AI research engineer working on large language models. He writes the Ahead of AI newsletter and is the author of Build a Large Language Model (From Scratch).
Now we know: The popular Ox Alpha LLM was GLM-5.3-Flash... Compared to GLM-5.2, this new GLM-5.3-Flash model uses: - a Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-heat Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers; - a scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B; - a DeepSeek V4-style mHC residual path with four parallel streams; - plus a native vision encoder (not shown). * "Super hybrid" because both KDA and MLA/DSA are "efficient" components. E.g., Kimi only uses KDA +… Quoting @Zai_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now… LLMs Related
19 August
Researcher and science writer. Co-founder and editor at Works in Progress magazine, formerly a researcher at Our World in Data, and author of the Scientific Discovery newsletter on science, global health and medical innovation.
It's unclear to me at the moment whether the view counter on my TED talk is broken or if my aunts and uncles have been frantically refreshing the page thousands of times each Related
4 June
JavaScript programmer; creator of the jQuery library, author of Secrets of the JavaScript Ninja, and long-time engineer at Khan Academy.
My pie-in-the-sky requests for @CloudflareDev to overhaul corporate web dev: * Fully hosted private Git + CI + "Pull Requests" * Playwright end-to-end test running and result collection * Storybook hosting + shapshot comparisons * Private NPM package registry hosting Related
21 December 2025
Economist; Professor of the Practice of Economic Policy at Harvard, nonresident senior fellow at the Peterson Institute, and Chairman of the US Council of Economic Advisers from 2013 to 2017.
Their words
But it also is exciting and refreshing to see it addressed with such originality and cutting-edge research.
Show the whole quote
fivebooks.com
29 October 2025
Game designer; wrote A Theory of Fun for Game Design and worked on Ultima Online and Star Wars Galaxies.
Stars Reach visual upgrades For months now, we have been working on a big overhaul to the visuals in Stars Reach. This has involved redoing most of the shaders in the entire game, and going back and retouching every asset. We started by redoing al…
Related
30 June 2025
Writes Don't Worry About the Vase, a near-daily account of what is happening in AI and what they think it means. Former Magic: the Gathering pro and trader; unusually willing to put a number on a belief and to grade their own past calls.
Their words
There is something highly refreshing about the way Kling offers his own consistent old school economic libertarian perspective, takes his time responding to anything, and generally tries to understand everything including developments in AI from that perspective.
Show the whole quote
thezvi.substack.com
21 March 2025
Writer and former McKinsey/BCG strategy consultant; author of The Pathless Path and Good Work, about leaving the default career track. Publishes the Boundless newsletter and podcast.
Their words
A refreshing counterargument to our culture's obsession with busyness, showing how deliberate rest is an essential component of the good life.
Show the whole quote
pathlesspath.com
31 December 2024
Machine learning engineer working on recommender systems, personalization and information retrieval. Previously worked on LLMs and LLM infrastructure at Mozilla.ai and on ML and recsys at Duo, Tumblr, Automattic and Comcast; wrote the 'What are embeddings?' text and ran the Normconf conference.
Their words
The Rattle Bag edited by Seamus Heaney and Ted Hughes - In an age of LLM-generated poetry and machine learning curation (even coming from yours truly sometimes), it's refreshing to have real people who love poetry curate it and show you what you need to read to touch grass.
Show the whole quote
vickiboykis.com
9 June 2022
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024, latterly leading its safety systems team.
Generalized Visual Language Models Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visua…
Related
Nothing matches. Show everything
What is a korrent?
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com .
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
Got it
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.
Got it