Artwork

Content provided by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka®. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka® or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.
Player FM - Aplicație Podcast
Treceți offline cu aplicația Player FM !

Next-Gen Data Modeling, Integrity, and Governance with YODA

55:55
 
Distribuie
 

Manage episode 357219000 series 2355972
Content provided by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka®. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka® or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

In this episode, Kris interviews Doron Porat, Director of Infrastructure at Yotpo, and Liran Yogev, Director of Engineering at ZipRecruiter (formerly at Yotpo), about their experiences and strategies in dealing with data modeling at scale.
Yotpo has a vast and active data lake, comprising thousands of datasets that are processed by different engines, primarily Apache Spark™. They wanted to provide users with self-service tools for generating and utilizing data with maximum flexibility, but encountered difficulties, including poor standardization, low data reusability, limited data lineage, and unreliable datasets.
The team realized that Yotpo's modeling layer, which defines the structure and relationships of the data, needed to be separated from the execution layer, which defines and processes operations on the data.
This separation would give programmers better visibility into data pipelines across all execution engines, storage methods, and formats, as well as more governance control for exploration and automation.
To address these issues, they developed YODA, an internal tool that combines excellent developer experience, DBT, Databricks, Airflow, Looker and more, with a strong CI/CD and orchestration layer.
Yotpo is a B2B, SaaS e-commerce marketing platform that provides businesses with the necessary tools for accurate customer analytics, remarketing, support messaging, and more.
ZipRecruiter is a job site that utilizes AI matching to help businesses find the right candidates for their open roles.
EPISODE LINKS

  continue reading

Capitole

1. Intro (00:00:00)

2. What is Yotpo? (00:02:29)

3. Building an ETL framework based on Spark (00:05:25)

4. What is Apache Spark? (00:10:18)

5. Decoupling the data model (00:15:40)

6. Using data mesh principles (00:18:51)

7. How to address different data personas (00:22:24)

8. What is the "shift left" movement? (00:26:35)

9. How can organizations change the way they treat their data? (00:28:47)

10. Use-cases for tooling and documenting data sets (00:31:01)

11. Schema vs. schema-less (00:32:07)

12. What is YODA? (00:40:07)

13. Takeaways from the conversation with Doron and Liran (00:48:35)

14. It's a wrap! (00:52:45)

265 episoade

Artwork
iconDistribuie
 
Manage episode 357219000 series 2355972
Content provided by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka®. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Confluent, founded by the original creators of Apache Kafka® and Founded by the original creators of Apache Kafka® or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://ro.player.fm/legal.

In this episode, Kris interviews Doron Porat, Director of Infrastructure at Yotpo, and Liran Yogev, Director of Engineering at ZipRecruiter (formerly at Yotpo), about their experiences and strategies in dealing with data modeling at scale.
Yotpo has a vast and active data lake, comprising thousands of datasets that are processed by different engines, primarily Apache Spark™. They wanted to provide users with self-service tools for generating and utilizing data with maximum flexibility, but encountered difficulties, including poor standardization, low data reusability, limited data lineage, and unreliable datasets.
The team realized that Yotpo's modeling layer, which defines the structure and relationships of the data, needed to be separated from the execution layer, which defines and processes operations on the data.
This separation would give programmers better visibility into data pipelines across all execution engines, storage methods, and formats, as well as more governance control for exploration and automation.
To address these issues, they developed YODA, an internal tool that combines excellent developer experience, DBT, Databricks, Airflow, Looker and more, with a strong CI/CD and orchestration layer.
Yotpo is a B2B, SaaS e-commerce marketing platform that provides businesses with the necessary tools for accurate customer analytics, remarketing, support messaging, and more.
ZipRecruiter is a job site that utilizes AI matching to help businesses find the right candidates for their open roles.
EPISODE LINKS

  continue reading

Capitole

1. Intro (00:00:00)

2. What is Yotpo? (00:02:29)

3. Building an ETL framework based on Spark (00:05:25)

4. What is Apache Spark? (00:10:18)

5. Decoupling the data model (00:15:40)

6. Using data mesh principles (00:18:51)

7. How to address different data personas (00:22:24)

8. What is the "shift left" movement? (00:26:35)

9. How can organizations change the way they treat their data? (00:28:47)

10. Use-cases for tooling and documenting data sets (00:31:01)

11. Schema vs. schema-less (00:32:07)

12. What is YODA? (00:40:07)

13. Takeaways from the conversation with Doron and Liran (00:48:35)

14. It's a wrap! (00:52:45)

265 episoade

Toate episoadele

×
 
Loading …

Bun venit la Player FM!

Player FM scanează web-ul pentru podcast-uri de înaltă calitate pentru a vă putea bucura acum. Este cea mai bună aplicație pentru podcast și funcționează pe Android, iPhone și pe web. Înscrieți-vă pentru a sincroniza abonamentele pe toate dispozitivele.

 

Ghid rapid de referință