Search
120 results
Said and published
-
testit — A simple package for testing R packages
-
Unit tests have become just more source code to maintain — only end-to-end tests can be trusted to guard real behavior.
unit tests are just source code at this point e2e test is the only real kind of test we can trust to guard actual behavior
-
awesome-website-testing-tools — Resource of web-based testing and validation tools
-
Belly fat is killing your testosterone
-
insta — A snapshot testing library for rust
-
Review, testing and QA is the new bottleneck in software engineering.
the new bottleneck in software engineering: review, testing and QA
-
Development should never be driven by tests: how much to test is an engineering decision to be costed project by project, not a methodology to adopt by default.
I would say the part that I don't like about test-driven development is the testdriven part. I don't think development should ever be driven by tests.
-
inputmodes.com — Testing inputmode
-
lablaudo-bot — vibe coded telegram bot to check blood test results
-
BrowserStack — BrowserStack
BrowserStack Cross-device testing
-
One of the kids got a surprisingly bad grade on a spelling pre-test, so I gave Codex a picture of the test and told it to build a practice app. Much improved!
-
How well do agents use test/verification techniques?
-
How well do agents use test/verification techniques?
-
How well do agents use test/verification techniques?
-
Big TestFlight update for Unforgetful:
-
AI may finally allow testing whether or where people make more money by being good.
AI may finally allow us to test whether (or more precisely where) you make more money by being good.
-
Completed my first FTP test yesterday with @GoZwift 💀
-
Test-time scaling has gained a third axis of latent space reasoning iterations in looped transformers.
Seems like test time scaling has gained a 3rd axis: latent space reasoning iterations in looped transformers.
-
A Clear and Classic Big Test of U.S. Power
-
Konbini Taste Test! Ready... set... EAT! 🍱 🍙 🍡
-
csi — Tool to test terminal csi commands
-
Just testing my font in a few random words &
-
Just testing my font in a few random words &
-
Testimony Data Project
-
Everyone Is Talking About Male Testosterone
-
Want to help me shape Unforgetful's direction? Join the TestFlight here: https://testflight.apple.com/join/Ex8upS2T Limited to the first 200 people. Actual users only, please. Thanks!
-
go-test-bug-repro
-
Uh oh, sounds like another agent escaped its testing environment!
-
Performance test results are highly dependent on the specific machine, library versions, and other components used, making them hard to compare over time.
Performance tests are highly specific and dependent on the exact machine it runs on, the exact third party libraries and their versions that are used, the other components involved in the tests, such as the servers, and…
-
I Delete Tests Every Night (On Purpose)
-
To whoever added several more steps to the task of releasing a TestFlight build to a group of testers in App Store Connect: 👎 👎 👎
-
…the Playbit runtime for Linux and it's surprisingly difficult to test various desktop configurations. We need a real GPU with a display connected, for testing dawn/webgpu, vsync, display sleep, etc. We use…
-
Twelve HIFi Amps Tested and the Winner Is....
-
Apple seriously added search to the TestFlight app, *and* removed sorting options, with the default now being alphabetical sorting only…? WTF? This makes TestFlight unusable for me.
-
Terrific testimony, worth reading.
-
Unit testing frameworks arrived so late because testing was a status divide, and testers had every incentive to keep a separate tool and language of their own.
So beforehand because of this social divide between programmers and testers. There was a lot of incentive for the testers to have their own language. This is my tool. I know how to run it. I'm going to run it.
-
At three billion users you cannot launch without testing and cannot test without it becoming news, so the communications plan comes before the launch decision.
you can't you can't you can't launch something to three billion people and not test it first, but you can't test something at our scale and not expect people to cover it
-
typeland — yet another typing test website (wip)
-
Aaron Newman: A Breakthrough Blood Test for Cancer
-
Nobody should buy a direct-to-consumer biological age test today: they are sold expensively without standardisation or demonstrated reliability.
I would suggest not buying any of these tests.
-
hegel-clj — Clojure bindings for Hegel, a property-based testing system.
-
Inform6-Testing — Regression testing for the Inform 6 compiler
-
fair-testing
-
The God Test
-
I tested EVERY single IP KVM
-
Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island:
-
The safety frameworks the labs publish do not account for how much test-time compute a dangerous-capability evaluation was given.
the preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
-
Helo — Beyond Code
For email testing, I love using Helo.
-
Testing Library
Testing Library - A great testing library for anything that interacts with the DOM.
-
Unit testing's only job should be to flag when something unexpected happens.
I still believe that unit testing should be nothing more than "tell me if something unexpected happened."
-
testing Vue components in the browser
-
Intelligence Testy
-
test-do-rpc-fetch-ordering
-
new Horse Race Tests premiering in cursor camp everyday this week! @snakesandrews
-
Assumptions weaken properties
-
Standardised testing is a good idea badly executed: the tests in common use are the problem, not testing itself.
Standardized testing is glorious, but many standardized tests royally suck.
-
Good things happen when you do not test implementation details: tests written with the React and Vue testing libraries look almost identical.
Did you know, the tests you write with react-testing-library look almost identical to the tests you write with vue-testing-library. Magic things happen when you don't test implementation details. :)
-
random-python-app — Ignore, this is a test of a skill
-
cf-allow-test
-
go-testutil — Go test helper
-
Go Testing By Example
-
express-unit-test — Simple Express unit test example.
-
vue-testing-library-sample — 🚀 A Vue.js project to test Jest and Testing-library. Data come from Star Wars API.
-
js-unit-testing-examples — 🤓 JavaScript Unit Testing Examples
-
Test coverage is a tool for finding untested code and says nothing about how good the tests are.
Test coverage is of little use as a numeric statement of how good your tests are.
-
Vitest
Vitest - A great testing framework.
-
Improving Your Workflow with JavaScript Testing: Best Practices
-
test-zod
-
hacker-test-history — Let's explain all the hacker test questions!
-
test-deleteme
-
HTTP3-test — Documentation for early HTTP/3 testing (with curl and more)
-
Writing code with a coding agent and no tests at all is indefensible, because the old objection to testing — that it is extra work you then have to maintain — no longer applies.
I think I see people who are writing code with coding agents and they're not writing any tests at all. That's a terrible idea.
-
telegram-test-bot — A test Telegram bot using serverless functions with Vercel for deployment
-
Write tests. Not too many. Mostly integration.
Write tests. Not too many. Mostly integration.
-
SAE J3018 for operational safety
-
In an agent-written codebase the agent-written tests are equally untrustworthy; manually using the product is the only reliable measure of whether it works.
Worse, you realize that the gazillions of unit, snapshot, and e2e tests you had your clankers write are equally untrustworthy. The only thing that's still a reliable measure of "does this work" is manually testing the…
-
Playwright — Microsoft
Playwright - I use this for E2E testing.
-
test
-
testing
-
Most regulatory delay is dead calendar time that could be cut without changing a single safety test.
There is a massive scope to reduce calendar time without even altering testing protocols.
-
Unit tests alone give little confidence in a program, because the mistakes people actually make are not the conveniently unit-testable ones.
But unit tests don't give you much confidence in your code.
-
Conventional testing is not sufficient: crude purpose-built checking tools of a few thousand lines find bugs that testing does not.
These studies also hammer home the point that conventional testing isn't sufficient.
-
sphinx-github-action-test — A quick repo for testing compiling a sphinx doc and syncing it with S3
-
workspace-testing
-
Buggy code usually causes outright test failures rather than being silently exercised inside tests that still pass.
a passing test can still execute buggy code if the bug is data-dependent, or if the test is not sensitive to the specific mistake in the code. But a lot of the time, buggy code only triggers failures.
-
dotfiles — testing dotfiles
-
go-with-test
-
Observatory — Testing Observable
-
testing_with_zeb
-
simplelab — SimpleLab
I use simplelab to test the water.
-
Updating to .NET 8, updating to IHostBuilder, and running Playwright Tests within NUnit headless or headed on any OS
-
Tests written by a model in the same context as the change are what find the bugs; the automated tests left behind are the lesser product.
Bigger changes always get tests. Automated ones usually aren’t great, but the model almost always finds issues when you ask it to write tests IN THE SAME CONTEXT. Context is precious, don’t waste it.
-
Don't Trip[wire] Yourself: Testing Error Recovery in Zig
-
A/B testing cannot touch anything that actually decides a company: not strategy, not vision, not insight.
AB testing doesn't work very well and it doesn't work on most things. It won't work on strategy or vision or insights like nothing actually important to the success of the company. You don't AB test whether Uber is a…
-
Grading flattens whatever is taught into what can be tested, and teaches students to ignore everything else.
Evaluation forces me to flatten everything I teach into something that can be tested, and it encourages students to ignore everything that isn’t on the test.
-
Almost every software project has cheap testing wins left on the table: basic random testing, static analysis and fault injection pay for themselves the first time they are run.
Almost every software project I've seen has a lot of low hanging testing fruit.
-
A performance test without an explanation of its limiting factor should be treated as unanalyzed and possibly bogus.
Any performance test should be accompanied by an explanation of the limiting factor, since no explanation will reveal the test wasn't analyzed and the result may be bogus.
-
How much to test is an economic calculation a team should actually run, not a matter of principle.
The point isn't that you should definitely write more tests, it's that you should definitely do the math to see if you should write more tests.
-
gollum-demo — Gollum test repo
-
Working Effectively with Legacy Code — Michael Feathers
Legacy code has no test and is not written to be testable. Touching it this sphaghetti code somewhere breaks the system. And to refactor safely, we'd need to have tests first... but where to start? This book gives…
-
Total test coverage is a warning sign rather than an achievement: it smells of tests written for the number.
I would be suspicious of anything like 100% - it would smell of someone writing tests to make the coverage numbers happy, but not thinking about what they are doing.
-
Write tests. Not too many. Mostly integration.
"Write tests. Not too many. Mostly integration." - @rauchg Here's my talk from @assertjs two weeks ago dissecting that statement.
-
github-actions-test — playing with GiHub Actions
-
Making test coverage a target destroys it, because high coverage numbers are easy to reach with worthless tests.
If you make a certain level of coverage a target, people will try to attain it. The trouble is that high coverage numbers are too easy to reach with low quality testing.
-
JSCheck — A random property testing tool for JavaScript
-
AB testing some Mac & Cheese recipes right now
-
Undetected AI Exam Answers
-
math_full_minus_math500 — The MATH training dataset without MATH-500 test portion
-
10 Tips for Writing Better Tests
-
Hoppscotch
Hoppscotch for testing APIs, after Insomnia turned to 💩
-
LLMs will push software development toward more specialized code, fewer generalized packages, and more readable tests.
So I foresee a world with far more specialized code, with fewer generalized packages, and more readable tests.
-
Code that is complex or clever can pass its tests while still being bad code.
A lot of complex, "look at this clever trick", overly-abstracted, unreadable code works and makes the tests pass.
-
The New Business Road Test — John Mullins
These books are more specific and rigorous than your standard Malcolm Gladwell fare; things like Getting to Plan B, The New Business Road Test, and Business Model Generation.
-
Bad admissions tests survive on herding: no programme wants to be the one asking for a different exam, even a far better one.
So a lot of the problem is mere herding: If every other top engineering program is using the GRE, it’s too weird to require a different test even if the alternative test is a far superior selection mechanism.
-
Most distributed-systems outages would have been prevented by slightly more comprehensive ordinary testing, not by better tools.
In fact, the vast majority of distributed systems outages could have been prevented by slightly-more-comprehensive testing.
-
The one thing no science-fiction writer dared predict is that the Turing test would be passed and nobody would notice.
And you know what no sci-fi author dared to predict is that the Turing tests just passes by.
-
Testing shades of a colour is an early-2000s tactic that no longer drives anything.
For the love of God, a blue is a blue is a blue. As long as it's accessible and as long as it's bright enough, off you go. You do not need to test the shades of blue, but test it against green or so on.
-
The more your tests resemble the way your software is used, the more confidence they can give you.
The more your tests resemble the way your software is used, the more confidence they can give you.
-
The leading chatbots are undifferentiated: in a blind test most people could not tell one model's output from another's.
Like it seems to me right now you could do like a double blind test of the same prompt given to Grock Claude Gemini um Mistral Deep Seek. Do a double blind test. I bet most people wouldn't be able to tell which is which.
-
If a test cannot collect its sample size within a month, it should not be run at all.
My rule of thumb, if we cannot collect the sample size in a month, we shouldn't test it. Period.