Jimmy Ba
Assistant Professor of Computer Science at the University of Toronto; co-creator, with Diederik Kingma, of the Adam optimizer for training neural networks.
Jimmy Ba did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
Beliefs
Korrents What they believe 4 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments.
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train.
The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
The same optimizer should handle objectives that move under it and gradients that are noisy or sparse.
The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
Show 1 more
An optimizer's hyper-parameters should mean something a practitioner can reason about, and should rarely need tuning.
The hyper-parameters have intuitive interpretations and typically require little tuning.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
Beliefs others hold too
Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. 2 hold this
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. 2 hold this
The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
The same optimizer should handle objectives that move under it and gradients that are noisy or sparse. 2 hold this
The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients.
Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv) Said 22 Dec 2014
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.