Illustrating Reinforcement Learning From Human Feedback (RLHF) AI Safety Fundamentals: Alignment podcast

Artwork

Tech Society Philosophy Blue Dot Impact

Sisällön tarjoaa BlueDot Impact. BlueDot Impact tai sen podcast-alustan kumppani lataa ja toimittaa kaiken podcast-sisällön, mukaan lukien jaksot, grafiikat ja podcast-kuvaukset. Jos uskot jonkun käyttävän tekijänoikeudella suojattua teostasi ilman lupaasi, voit seurata tässä https://fi.player.fm/legal kuvattua prosessia.

AI Safety Fundamentals: Alignment « »
Illustrating Reinforcement Learning from Human Feedback (RLHF)

7M ago 22:32

Jaa

MP3•Jakson koti

Arkistoidut sarjat ("Toimeton syöte" status)

When? This feed was archived on February 21, 2025 21:08 (5d ago). Last successful fetch was on January 02, 2025 12:05 (2M ago)

Why? Toimeton syöte status. Palvelimemme eivät voineet hakea voimassa olevaa podcast-syötettä tietyltä ajanjaksolta.

What now? You might be able to find a more up-to-date version using the search function. This series will no longer be checked for updates. If you believe this to be in error, please check if the publisher's feed link below is valid and contact support to request the feed be restored or if you have any other concerns about this.

Sisällön tarjoaa BlueDot Impact. BlueDot Impact tai sen podcast-alustan kumppani lataa ja toimittaa kaiken podcast-sisällön, mukaan lukien jaksot, grafiikat ja podcast-kuvaukset. Jos uskot jonkun käyttävän tekijänoikeudella suojattua teostasi ilman lupaasi, voit seurata tässä https://fi.player.fm/legal kuvattua prosessia.

This more technical article explains the motivations for a system like RLHF, and adds additional concrete details as to how the RLHF approach is applied to neural networks.

While reading, consider which parts of the technical implementation correspond to the 'values coach' and 'coherence coach' from the previous video.

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

… continue reading

Luvut

1. Illustrating Reinforcement Learning from Human Feedback (RLHF) (00:00:00)

2. RLHF: Let’s take it step by step (00:03:16)

3. Pretraining language models (00:03:51)

4. Reward model training (00:05:46)

5. Fine-tuning with RL (00:09:26)

6. Open-source tools for RLHF (00:16:10)

7. What’s next for RLHF? (00:18:20)

8. Further reading (00:21:17)

85 jaksoa

#Tech #Society #Philosophy #Blue Dot Impact

Artwork

Illustrating Reinforcement Learning from Human Feedback (RLHF)

AI Safety Fundamentals: Alignment

published 7M ago

Jaa

MP3•Jakson koti

Arkistoidut sarjat ("Toimeton syöte" status)

When? This feed was archived on February 21, 2025 21:08 (5d ago). Last successful fetch was on January 02, 2025 12:05 (2M ago)

Why? Toimeton syöte status. Palvelimemme eivät voineet hakea voimassa olevaa podcast-syötettä tietyltä ajanjaksolta.

What now? You might be able to find a more up-to-date version using the search function. This series will no longer be checked for updates. If you believe this to be in error, please check if the publisher's feed link below is valid and contact support to request the feed be restored or if you have any other concerns about this.

Sisällön tarjoaa BlueDot Impact. BlueDot Impact tai sen podcast-alustan kumppani lataa ja toimittaa kaiken podcast-sisällön, mukaan lukien jaksot, grafiikat ja podcast-kuvaukset. Jos uskot jonkun käyttävän tekijänoikeudella suojattua teostasi ilman lupaasi, voit seurata tässä https://fi.player.fm/legal kuvattua prosessia.

This more technical article explains the motivations for a system like RLHF, and adds additional concrete details as to how the RLHF approach is applied to neural networks.

While reading, consider which parts of the technical implementation correspond to the 'values coach' and 'coherence coach' from the previous video.

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

… continue reading

Luvut

1. Illustrating Reinforcement Learning from Human Feedback (RLHF) (00:00:00)

2. RLHF: Let’s take it step by step (00:03:16)

3. Pretraining language models (00:03:51)

4. Reward model training (00:05:46)

5. Fine-tuning with RL (00:09:26)

6. Open-source tools for RLHF (00:16:10)

7. What’s next for RLHF? (00:18:20)

8. Further reading (00:21:17)

85 jaksoa

#Tech #Society #Philosophy #Blue Dot Impact

ทุกตอน

×

Tervetuloa Player FM:n!

Player FM skannaa verkkoa löytääkseen korkealaatuisia podcasteja, joista voit nauttia juuri nyt. Se on paras podcast-sovellus ja toimii Androidilla, iPhonela, ja verkossa. Rekisteröidy sykronoidaksesi tilaukset laitteiden välillä.

Kuuntele yli 500 aihetta

Pikakäyttöopas

Suosituimmat podcastit

Lindgren & Sihvonen

Urheilun ääni

Nordea Markets Insights FI

Kasper ja Mikko - Suomen suosituin podcast

Kolme miestä ja elokuvasauva

Pyöreä pöytä

Uutisraportti podcast

Stressivapaa johtaja | Näkökulmia henkilökohtaiseen kasvuun

Vinkistä vihiä

Apua/UKK | Päivitä | Mainostaa

Taide|Liike-elämä|Komedia|Talous|Viihde|Uutiset|Politiikka|Uskonto

Tiede|Jalkapallo|Urheilu|Tarinankerronta|Teknologia|True crime

Tekijänoikeudet 2025 | Sivukartta | Tietosuojakäytäntö | Käyttöehdot | | Tekijänoikeus

Kuuntele tämä ohjelma tutkiessasi