Diederik P. Kingma
Research scientist at Anthropic; co-creator of the Variational Autoencoder and, with Jimmy Ba, of the Adam optimizer, two of the most widely used tools in deep learning.
Diederik P. Kingma did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
Mastodon @dpkingma@sigmoid.social The same, on Mastodon.
His secondary account on the ML-focused instance sigmoid.social.
Recent
- “Adam can converge without any modification on update rules” https://arxiv.org/abs/2208.09632 Proves that (vanilla) Adam is theoretically justified without any modification. Presented at NeurIPS'22. By Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, Zhi-Quan Luo. Also provides suggestions for tuning hyperparameters beta1 and beta2. 3 Dec 2022
- Pros and Cons of SF, according to chat.openai.com. 1 Dec 2022
- A poster with eight tablets for animations. Wild. Full-sized rollable OLED screens as posters in the near future? #NeurIPS #NeurIPS22 30 Nov 2022
Show 4 more
- Somebody should make bottom-loading kitchen ovens. Could be wall- or ceiling-mounted. Much more energy efficient, and your face wouldn't get blasted every time you open the oven door. Also, kid safe. 28 Nov 2022
- I'll be at NeurIPS in New Orleans from Tuesday to Saturday. Who is going? Let me know if you have a poster presentation I should definitely hit up! Looking forward to catching up! 27 Nov 2022
- This might be blasphemy over here, but to support longer/deeper discussions on Mastodon, I think it'd be nice if Mastodon would implement options for non-chronological lists of posts, e.g. similar to reddit's default 'hot" view. Could be as simple as sorting threads based on most recent reply. This could be opt-in and can co-exist with the current chronological view. 9 Nov 2022
- Just migrated to sigmoid.social, an AI-focussed Mastodon server. Thanks for setting it up @thegradient! Migration was a breeze. 5 Nov 2022
Link verified 9 Sept 2026. Recent items update automatically from the channel.
Beliefs
Korrents What they believe 4 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments.
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train.
The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
The same optimizer should handle objectives that move under it and gradients that are noisy or sparse.
The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
Show 1 more
An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning.
The hyper-parameters have intuitive interpretations and typically require little tuning.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
Beliefs others hold too
Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. 2 hold this
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. 2 hold this
The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
The same optimizer should handle objectives that move under it and gradients that are noisy or sparse. 2 hold this
The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.