Artwork

Content provided by BlueDot Impact. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by BlueDot Impact or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.
Player FM - Aplicație Podcast
Treceți offline cu aplicația Player FM !

AI Control: Improving Safety Despite Intentional Subversion

20:51
 
Distribuie
 

Serii arhivate ("Sursă inactivă" status)

When? This feed was archived on February 21, 2025 21:08 (5d ago). Last successful fetch was on January 02, 2025 12:05 (2M ago)

Why? Sursă inactivă status. Servele noastre nu au putut să preia o sursă valida de podcast pentru o perioadă îndelungată.

What now? You might be able to find a more up-to-date version using the search function. This series will no longer be checked for updates. If you believe this to be in error, please check if the publisher's feed link below is valid and contact support to request the feed be restored or if you have any other concerns about this.

Manage episode 424744791 series 3498845
Content provided by BlueDot Impact. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by BlueDot Impact or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post:

  • We summarize the paper;
  • We compare our methodology to the methodology of other safety papers.

Source:
https://www.alignmentforum.org/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion
Narrated for AI Safety Fundamentals by Perrin Walker

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

  continue reading

Capitole

1. AI Control: Improving Safety Despite Intentional Subversion (00:00:00)

2. Paper summary (00:02:41)

3. Setup (00:02:43)

4. Evaluation methodology (00:04:59)

5. Results (00:06:25)

6. Relationship to other work (00:10:51)

7. Future work (00:17:50)

85 episoade

Artwork
iconDistribuie
 

Serii arhivate ("Sursă inactivă" status)

When? This feed was archived on February 21, 2025 21:08 (5d ago). Last successful fetch was on January 02, 2025 12:05 (2M ago)

Why? Sursă inactivă status. Servele noastre nu au putut să preia o sursă valida de podcast pentru o perioadă îndelungată.

What now? You might be able to find a more up-to-date version using the search function. This series will no longer be checked for updates. If you believe this to be in error, please check if the publisher's feed link below is valid and contact support to request the feed be restored or if you have any other concerns about this.

Manage episode 424744791 series 3498845
Content provided by BlueDot Impact. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by BlueDot Impact or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post:

  • We summarize the paper;
  • We compare our methodology to the methodology of other safety papers.

Source:
https://www.alignmentforum.org/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion
Narrated for AI Safety Fundamentals by Perrin Walker

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

  continue reading

Capitole

1. AI Control: Improving Safety Despite Intentional Subversion (00:00:00)

2. Paper summary (00:02:41)

3. Setup (00:02:43)

4. Evaluation methodology (00:04:59)

5. Results (00:06:25)

6. Relationship to other work (00:10:51)

7. Future work (00:17:50)

85 episoade

Toate episoadele

×
 
Loading …

Bun venit la Player FM!

Player FM scanează web-ul pentru podcast-uri de înaltă calitate pentru a vă putea bucura acum. Este cea mai bună aplicație pentru podcast și funcționează pe Android, iPhone și pe web. Înscrieți-vă pentru a sincroniza abonamentele pe toate dispozitivele.

 

Ghid rapid de referință

Listen to this show while you explore
Play