These mocha-ordering fellow humans must fall into one of these categories: They don't know that it's rude They don't care that it's rude They disagree that it's rude
AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at.
But from their perspective, they've just been trained for millions of subjective years to do as well as they possibly can on these evals. In many cases, the only way in which they've been able to perform well on that training is explicitly by cheating, right?
born free technologies are ones where there aren't regulations on the books. Things of like the internet when it just started. And those are the types of technologies you have to preserve.
So if you design call it loop, call it graph, call it workflow, it's kind of all the same. If you design something that gets a trigger or gets an input and does something for you and maybe there's a decision in the way, boom, there's your graph.
Anyone using GitHub knows that it isn't cut out for agents. The workflow, the review process, the merge queue and many things are made for the previous era.
In addition to reliability issues, it often engenders a mind-numbing workflow and an environment where junior developers will never acquire the expertise to become senior developers capable of designing complex systems.
I think evals, they outlive the harness a little bit, but not by that much. Like an eval might live for maybe one, two, three model generations, but nowadays the you know, we're on the exponential. The model is improving so quickly, very often we just saturate the eval, and then we have to throw it away, and we have to come up with a new eval.
Which means we need to have a flexibility in the tools that we provide and the types of partnerships we have, and to really have a creator enablement view rather than a prescriptive that we only do this one way.
It does feel like we have to be more more explicit about the types of people and talent that tend to thrive at Netflix versus other companies like some of the frontier labs.
I get really nervous about having different design languages or different types of user interactions and shipping Frankensteins, basically. So, designers need to then be the people we're hiring again for design systems thinking.
They get women to sleep with them, but they attract both men and women and cause them to surrender their better judgment and make bad choices in a variety of contexts, including the voting booth.
The thing that makes Kubernetes powerful, there's a data model. We gave infrastructure a type system. So instead of imperative shell scripts, you finally had types.
And the reason they came, I think, is absolutely because of the the better tooling. And I think we were totally right there that like adding a an erasable type system and then using that to enable great tooling is really where the pro where the programmer productivity boost is realized.
The thing that makes it interesting, I think, and and unlike pretty much any other programming language is the gradual typing. This This notion that you can have types, but you don't have to have types.
because if you were to force AI to write a type annotation on everything, then it would probably get it wrong more often because now it has to keep track of all these types and and it and it has to just repeat itself over and over and over, right? And so, types are important where there's no context.
So, if we're checking 99% instead of 100%, well, heck, that's better than the 0% that JavaScript checked, right? And it gives you like language features that no other languages can provide because they can't get to 100%.
the companies that are very successful actually have both types of organizations inside their company. And that the leaders of the organization are the ones who are responsible for creating a healthy, functioning relationship between the two types of organizations.
there's sort of there's two types of work that tends to be involved in any kind of creative project. There's routine stuff and there you just want to avoid procrastination. You just want to like, you know, how do I get good at this or how do I outsource it and how do I do it as rapidly as possible? Um and just avoid you know, like getting into a situation where you're prolonging it. Um, and then there's high variance stuff where you actually you need to be willing to to you know take a lot of time.
So, I I think there's a lot of things we still try to imagine. What would be an AI-driven world look like, right? I think people are trying to like retrofit what already exists to fit what they think is a new workflow.
So, I think that the whole workflow of like reviewing code is very outdated. Like, I don't think the I think that the senior member, instead of like giving feedback on the code, they should be giving feedback on like how you give instruction to AI to produce better.
you know, right now we're going through uh an in um an cognitive version of the Copernican revolution where we used to think that human intelligence is the center of the universe. And now we're actually seeing that there's there's very different types of intelligence um that that that are out there uh with very different strengths and weaknesses.
We all, to some degree, lack the self-awareness of how we tick. So we’re all different types of gamers, but if you ask me to describe the type of gamer I am, I might actually be giving more of a picture of the type of gamer I wish I was or the type of gamer I want you to think I am versus the type of gamer I actually am.
Like models are good at different types of coding. Models have different styles. Like I think I think these things are actually, you know, quite different from each other. And so expect more differentiation than you see in in um cloud.
Immediately cease trying to perform meaningful work via a chatbot (e.g. ChatGPT, Gemini on the web, etc.). Chatbots have real value and are a daily part of my AI workflow, but their utility in coding is highly limited because you're mostly hoping they come up with the right results based on their prior training, and correcting them involves a human (you) to tell them they're wrong repeatedly.
the big unknown to get to whole body reversible is the brain. Um it's unclear like like the brain can withstand a lot of change and does withstand a lot of different types of damage or or change with age for example.
The US has a lot to learn from other countries for how to catch up to the efficient frontier, especially for certain types of construction (ie: transit) and in certain places in the country (ie: expensive coastal metros).
it's kind of like databases right it's always the thing it's like hey can one database be the one that just is used everywhere except it's not uh there are multiple types of databases that are getting deployed uh for different use cases.
even if the tech is diffusing fast uh this time around for true economic growth to appear it has to sort of diffuse to a point where the work the work artifact and the workflow has to change and so that's kind of one place where I think uh the change management required for a corporation to truly change I think is something we shouldn't discount
I would say though that something that is kind of annoying to me is that we haven't yet figured out the bridging from the tinkering to the workflow quite as seamlessly as I would like.
The overall workflow is intuitive, especially for those new to formal evaluation processes. The UI guides you through creating datasets, running experiments, and annotating results.
Decades of grading data; standardized test scores; cross-sectional, longitudinal, observational, and experimental studies; along with many other types of ancillary and convergent evidence, ultimately tell the same story: education can raise the absolute performance of most students modestly, but it almost never meaningfully reshuffles the relative distribution of ability and achievement
I'm much more comfortable with the fox paradigm. Yeah. So yeah, I like looking for analogies, narratives. I spend a lot of time… If there's a result, I see it in one field, and I like the result, it's a cool result, but I don't like the proof, it uses types of mathematics that I'm not super familiar with, I often try to re-prove it myself using the tools that I favor.
For most of my investing life, Vanguard was THE one-stop shop for index funds of all types. They have the lowest expense ratio and the utmost respect for their customers.
And the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
You might be skeptical of using synthetic data. After all, it’s not real data, so how can it be a good proxy? In my experience, it works surprisingly well. Some of my favorite AI products, like Hex use synthetic data to power their evals
One of the fascinating things that DNA is showing us, which actually blood types were showing us way before that, is that the oldest people in the Americas are in South America, the ones that got separated early and didn’t mix their DNA, like the people in the Amazon.
One thing that drives me crazy about the React ecosystem, and more specifically "tech influencers" and "thought leaders" in the space, is the infantilization of the developers using and working in it. I’m tired of reading takes like TypeScript generics, mapped types, etc should only be needed and used by library authors for most use cases. Or today’s discourse; Don’t use `useCallback’, ‘useMemo’, and React.memo. If you’re building anything beyond a simple CRUD app or a an incredibly focused app with few features, you _will_ need these features.
the approach that I use is called Schema Therapy, and Schema Therapy is designed to work with very complex personality types, personality disorders, and even those that are happening along a spectrum
It's basically the Vim of audio editing software, so I wouldn't recommend it if you don't have the time and energy to heavily invest in customizing it for your workflow, but I can move like lightning in this thing.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.