The Most Important Algorithm in Machine Learning

What is mechanistic interpretability? Neel Nanda explains.

When Computers Write Proofs, What's the Point of Mathematicians?

🔴Live โหนกระแส ทหารเกณฑ์ช้ำรัก!!! บุกห้องพักเจอเมียถูสบู่ให้ชายชู้

Challenge: reach high⚽️🦶🏼#shorts #foryou #sports #soccer #challenge

เปิดสาเหตุ! "สิงคโปร์แอร์ไลน์" ลงจอดฉุกเฉิน ผู้โดยสารเสียชีวิต 1 เจ็บนับสิบ | ลุยชนข่าว | 21 พ.ค. 67

31 - Singular Learning Theory with Daniel Murfet

AXRP

มุมมอง 405

เพิ่มลงใน
- เพลย์ลิสต์ของฉัน
- ดูภายหลัง
แชร์

แชร์

ฝัง

ขนาดวิดีโอ:

แสดงแผงควบคุมโปรแกรมเล่น

เล่นอัตโนมัติ

เล่นใหม่

เผยแพร่เมื่อ 6 พ.ค. 2024
What's going on with deep learning? What sorts of models get learned, and what are the learning dynamics? Singular learning theory is a theory of Bayesian statistics broad enough in scope to encompass deep neural networks that may help answer these questions. In this episode, I speak with Daniel Murfet about this research program and what it tells us.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:26 - What is singular learning theory?
0:16:00 - Phase transitions
0:35:12 - Estimating the local learning coefficient
0:44:37 - Singular learning theory and generalization
1:00:39 - Singular learning theory vs other deep learning theory
1:17:06 - How singular learning theory hit AI alignment
1:33:12 - Payoffs of singular learning theory for AI alignment
1:59:36 - Does singular learning theory advance AI capabilities?
2:13:02 - Open problems in singular learning theory for AI alignment
2:20:53 - What is the singular fluctuation?
2:25:33 - How geometry relates to information
2:30:13 - Following Daniel Murfet's work
The transcript: axrp.net/episode/2024/05/07/e...
Daniel Murfet's twitter/X account: / danielmurfet
Developmental interpretability website: devinterp.com
Developmental interpretability TH-cam channel: / @devinterp
Main research discussed in this episode:
- Developmental Landscape of In-Context Learning: arxiv.org/abs/2402.02364
- Estimating the Local Learning Coefficient at Scale: arxiv.org/abs/2402.03698
- Simple versus Short: Higher-order degeneracy and error-correction: www.lesswrong.com/posts/nWRj6...
Other links:
- Algebraic Geometry and Statistical Learning Theory (the grey book): www.cambridge.org/core/books/...
- Mathematical Theory of Bayesian Statistics (the green book): www.routledge.com/Mathematica...
- In-context learning and induction heads: transformer-circuits.pub/2022...
- Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity: arxiv.org/abs/2106.15933
- A mathematical theory of semantic development in deep neural networks: www.pnas.org/doi/abs/10.1073/...
- Consideration on the Learning Efficiency Of Multiple-Layered Neural Networks with Linear Units: papers.ssrn.com/sol3/papers.c...
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks: arxiv.org/abs/1806.07572
- The Interpolating Information Criterion for Overparameterized Models: arxiv.org/abs/2307.07785
- Feature Learning in Infinite-Width Neural Networks: arxiv.org/abs/2011.14522
- A central AI alignment problem: capabilities generalization, and the sharp left turn: www.lesswrong.com/posts/GNhMP...
- Quantifying degeneracy in singular models via the learning coefficient: arxiv.org/abs/2308.12108
วิทยาศาสตร์และเทคโนโลยี

ความคิดเห็น • 2

@dizietz 10 วันที่ผ่านมา
This one was pretty technical for those of us that haven't read some of the foundational work for SLT. I had to stop and look up some specific details later, and still don't feel like I fully grasp what makes SLT different than other predictions about degeneracy and simple functions preference in terms of making predictions about nn behavior.
David's framing of fundamental structures in the data being more important across any training runs makes a lot of sense, I still don't grok how this helps with alignment. I suppose understanding stability of structure moves us closer, but both on something similar to interoperability but also on capabilities.
@nowithinkyouknowyourewrong8675 13 วันที่ผ่านมา
10ins in and I still don't get why it's interesting? like it's a math stats tool that he is still making that will enable us to do other things

ต่อไป

เล่นอัตโนมัติ

The Most Important Algorithm in Machine Learning

The Most Important Algorithm in Machine Learning

What is mechanistic interpretability? Neel Nanda explains.

What is mechanistic interpretability? Neel Nanda explains.

When Computers Write Proofs, What's the Point of Mathematicians?

When Computers Write Proofs, What's the Point of Mathematicians?

🔴Live โหนกระแส ทหารเกณฑ์ช้ำรัก!!! บุกห้องพักเจอเมียถูสบู่ให้ชายชู้

🔴Live โหนกระแส ทหารเกณฑ์ช้ำรัก!!! บุกห้องพักเจอเมียถูสบู่ให้ชายชู้

Challenge: reach high⚽️🦶🏼#shorts #foryou #sports #soccer #challenge

Challenge: reach high⚽️🦶🏼#shorts #foryou #sports #soccer #challenge

เปิดสาเหตุ! "สิงคโปร์แอร์ไลน์" ลงจอดฉุกเฉิน ผู้โดยสารเสียชีวิต 1 เจ็บนับสิบ | ลุยชนข่าว | 21 พ.ค. 67

เปิดสาเหตุ! "สิงคโปร์แอร์ไลน์" ลงจอดฉุกเฉิน ผู้โดยสารเสียชีวิต 1 เจ็บนับสิบ | ลุยชนข่าว | 21 พ.ค. 67

Live Score เชียร์บอล : ฟุตบอล รีโว่ คัพ 2023/24 บุรีรัมย์ ยูไนเต็ด vs เมืองทอง ยูไนเต็ด l REVO CUP

Live Score เชียร์บอล : ฟุตบอล รีโว่ คัพ 2023/24 บุรีรัมย์ ยูไนเต็ด vs เมืองทอง ยูไนเต็ด l REVO CUP

This is why Deep Learning is really weird.

This is why Deep Learning is really weird.

Building RAG at 5 different levels

Building RAG at 5 different levels

But what is a GPT? Visual intro to transformers | Chapter 5, Deep Learning

But what is a GPT? Visual intro to transformers | Chapter 5, Deep Learning

Machine Learning vs Deep Learning

Machine Learning vs Deep Learning

WE MUST ADD STRUCTURE TO DEEP LEARNING BECAUSE...

WE MUST ADD STRUCTURE TO DEEP LEARNING BECAUSE...

Gaussian Processes

Gaussian Processes

Building a neural network FROM SCRATCH (no Tensorflow/Pytorch, just numpy & math)

Building a neural network FROM SCRATCH (no Tensorflow/Pytorch, just numpy & math)

How The Most Useless Branch of Math Could Save Your Life

How The Most Useless Branch of Math Could Save Your Life

The Attention Mechanism in Large Language Models

The Attention Mechanism in Large Language Models

ใช้ Goodnotes 6 ยังไงให้คุ้มเรียนที่สุด, แอปจดโน้ตยอดฮิต ฟีเจอร์ที่ทำให้เรียนดีขึ้น✨ Peanut Butter

ใช้ Goodnotes 6 ยังไงให้คุ้มเรียนที่สุด, แอปจดโน้ตยอดฮิต ฟีเจอร์ที่ทำให้เรียนดีขึ้น✨ Peanut Butter

Power up all cell phones.

Power up all cell phones.

Super Editing Photos with Uipdate Gago Tools: Unlock the Top N Techniques || Editing Photo Update

Super Editing Photos with Uipdate Gago Tools: Unlock the Top N Techniques || Editing Photo Update

Microsoft Surface Copilot + PC Event: Everything Revealed in 13 Minutes

Microsoft Surface Copilot + PC Event: Everything Revealed in 13 Minutes

Wireless switch without wires Part 1

Wireless switch without wires Part 1

คอมข้างทาง I EP.2 (2.1) เดินตลาดริมถนนแล้วเจอเลย! งบ 3 ร้อย เปิดติดมั้ย? มาดู

คอมข้างทาง I EP.2 (2.1) เดินตลาดริมถนนแล้วเจอเลย! งบ 3 ร้อย เปิดติดมั้ย? มาดู

ประแจแหวนมัลติฟังก์ชัน

ประแจแหวนมัลติฟังก์ชัน

3D printed Nintendo Switch Game Carousel

3D printed Nintendo Switch Game Carousel