<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">

 <title>Pipeline Data Engineering Academy</title>
 <subtitle>Pipeline Academy was the world's first coding bootcamp focused on data engineering. Our curriculum, course material and writing on sustainable data craftsmanship stay online as a free archive.</subtitle>
 <link href="https://dataengineering.academy/atom.xml" rel="self"/>
 <link href="https://dataengineering.academy/"/>
 <updated>2026-08-27T17:15:46+00:00</updated>
 <id>https://dataengineering.academy/</id>
 <author>
   <name>Daniel Molnar</name>
   <email>info@dataengineering.academy</email>
 </author>

 
 <entry>
   <title>Pipeline Academy on Hiatus</title>
   <link href="https://dataengineering.academy/2022/06/22/see-you-soon.html"/>
   <updated>2022-06-22T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2022/06/22/see-you-soon.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;It’s time to share some important news with you: we’re taking time off to focus on our health and families, the launch of new data engineering cohorts is on hold until further notice.&lt;/p&gt;
&lt;h4 &gt;Health and family&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Running a bootstrapped company in times of repeated economic crises and data industry vibe shifts is a gift and a curse at the same time. No surprises here: it can be highly rewarding and joyful, but it can be exhausting and stressful as well. It’s June 2022; after two and a half years of pioneering the data engineering and data architecture education, and after celebrating multiple amazing cohorts of freshly minted data infrastructure experts, it’s time for Daniel and myself to turn the heat down and allow ourselves to recharge our batteries.&lt;/p&gt;
&lt;h4 &gt;Not a “Goodbye&quot;, but hopefully a “See you soon!”&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The current economic shift will continue to have major (and hopefully sobering) impact on the way data infrastructure and data teams are managed, so challenges in this space are abundant. If you’re looking to become the data architect who can make sense of the promises of the “modern data stack”, who can design, build and maintain data infra that just works, and who understands what really matters and what doesn’t, you’re still at the right spot.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;If you’d like to get in touch, feel free to &lt;a href=&quot;mailto:info@dataengineering.academy&quot;&gt;email us directly&lt;/a&gt;, or if you’d like to be notified when a new cohort launches, please sign up here.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;If you are looking for trainings for your organisation or partner up, just send us an &lt;a href=&quot;mailto:info@dataengineering.academy&quot;&gt;email&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We hope to see you soon,&lt;br&gt;Peter &amp;amp; Daniel&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Meet The Graduates: Guoda Paulikaite</title>
   <link href="https://dataengineering.academy/2022/05/02/meet-the-graduates-guoda-paulikaite.html"/>
   <updated>2022-05-02T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2022/05/02/meet-the-graduates-guoda-paulikaite.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we’ll share some of the stories that Daniel and I get to watch unfold at Pipeline Academy. Check out what our graduates have to say about the course, how they’ve tackled its challenges and what they are doing now with their new data engineering superpowers.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Peter: Can I ask you to please introduce yourself to the readers of Pipeline Academy’s blog? &lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Guoda: My name is Guoda, I&apos;m a data analyst. I am Lithuanian, I am also a mother. I&apos;ve learned sociology as my bachelor, which always will stay at my heart. And I&apos;m curious about how to structure all the uncertainty in the unstructured world that we live in - this is something that drives me. &lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/gouda_paulikaite_data_engineer_project_a_ventures.jpg&quot; alt=&quot;Guoda Paulikaite at Project A Ventures&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Guoda Paulikaite, Data Analyst and graduate of Pipeline Academy&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Just recently I joined a company that was represented by one of the guest speakers at the bootcamp, and that has really given a boost to my career in this field. Now that I&apos;m at &lt;a href=&quot;https://www.project-a.com&quot;&gt;Project A Ventures&lt;/a&gt;, I am absolutely having a blast, enjoying the work I do and, yeah... I mean… earlier I&apos;ve used to have a very clear answer to your question, but now I think I&apos;m at the beginning of something new where I&apos;m not only a data analyst or a business analyst. At the moment I am a data analyst that supports many startups and ventures in their requests for data analytics topics. I am an analyst who works on solving data problems to deliver actionable insights for a variety of business goals. But also shaping their strategies of how they&apos;re going to scale and what data needs they will have in data analytics, structures or architectures. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: So if I understand this correctly, you combine multiple skills in your current job: you are a data analyst, a data engineer, a data architect, and you&apos;re also involved in the business- and organisational development of various ventures, right?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: Yes, indeed. Actually, data analyst sounds like I&apos;m given data, I need to analyse it and that&apos;s pretty much it. But there is so much more that needs to happen before. And it seems that in the field of data there are many titles being created just to follow up with all new emerging required skills. So this is why I love my organisation for not worrying about this and just keep on focusing on the actual effort required. Yeah, leaving the content behind the titles flexible.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;So, you&apos;re totally right. I&apos;m part of managing projects, advising ventures on their high-level data architectures, investigating really specific technical solutions, thinking about processes, thinking about blockers. After all the goal is to analyse the data in a way that it&apos;s actionable, but it&apos;s just sort of the peak of the mountain. And I take a big part of what&apos;s underneath, working together with data engineers and different departments internally and externally to get to that that peak. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: What was your intention behind joining the data engineering course?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: Previously having worked in IT teams of large companies, my expectation for joining was to gain the confidence that I&apos;ve always felt some lack of, you know, coming from a different academic background, even if I was very interested in data and I was good at pushing IT projects forward. I had this idea that maybe I could just go ahead and learn something in a non-academic, non-traditional format with more recent technologies. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I really had a good idea from your marketing materials that it will be a course covering many different data fields. Actually, this was my highest wish and expectation that I will get a good understanding on all of them to be comfortable in dealing with different architectures, different situations, different projects without going into extreme depth or getting lost in details of just one field, but having that grand overview.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: How did three months of learning online feel for you? &lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: I would say, I would have preferred a physical classroom much more, but it went much better than I expected, and this was I think due to the group size that we’ve had (P: cohort of 9 individuals). The fear that I&apos;ve had about the online classroom was that I won&apos;t get enough feedback, or no one will notice my progress or when I&apos;m struggling. Or, you know, just miss out on having fun building relationships with your peers. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But dealing with this was so similar to how you go about it in the real world: your colleagues will not be around all the time, they will be somewhere else geographically and you have to figure out how to talk to them effectively. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: If you have to think back, it’s already been 10 months since you&apos;ve graduated. What are your most memorable moments? Highlights, lowlights, good things, bad things from the course. &lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: I do remember that the course was very challenging. It wasn&apos;t something where you sit back and relax and just get information stuffed in your head. It was many times so challenging that I was just asking myself how can I go forward. One of the highlights was how Daniel actually spotted this frustration and was able to manage it and bring me to take one step back together and say okay, this is all gonna be fine. It sounds quite straightforward and simple, but actually it&apos;s one of the things that you learn that the level of detail can be so overwhelming that it&apos;s hard to avoid getting lost. I think this is one of the skills that I&apos;ve learned.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The amount of information and amount of complexity makes it difficult for you to grasp how the world can work, how planes can fly and bank transactions can run if everything&apos;s so complex. Then I realised that no one actually has a simple answer. It&apos;s just a matter of experience and willingness. You know, not to get lost and just continue. Fail and stand up again. And when I realised that everyone is doing the same, like everyone, even the most experienced developers, they do the same thing. They just go on and search for a solution.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Did your perception of data engineering, or the role of a data engineer change throughout or after the course?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: It became much more down-to-earth than what I imagined before. I imagined that you really need to have a degree in computer science. I still consider that people who make that choice to study computer science get a really good advantage, a good basis. But after all, it&apos;s all about experience and curiosity — and if you have that, then the rewards are great.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Where do you think your career is headed right now? &lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: I&apos;m headed to a place where I&apos;m confident with the uncertainty with different architectures that I&apos;ve never seen before. There&apos;s always new emerging architectures, new technologies. Now I am able to have a conversation about them with confidence, and a structured way of approaching implementing them in projects, and managing the stakeholders along the way.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Confidence in uncertainty. I love this expression. Data engineers need a lot of this.&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;There are people out there who are considering embarking on this path of learning data engineering. Do you have an advice for them?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;G: I have so many different tips that I could share. Some of them are more specific, some are more about not losing motivation and persistence. But overall, I think my number one tip would be to try to find a problem that you think you can solve, to have one project that you can work on throughout the whole academy. A tangible, small problem you want to solve, not an abstract, academic one. That’s a great starting point.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Thank you, Guoda.&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Pipeline Academy Setting Trends at the EdTech Awards</title>
   <link href="https://dataengineering.academy/2022/04/14/pipeline-academy-at-the-edtech-awards.html"/>
   <updated>2022-04-14T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2022/04/14/pipeline-academy-at-the-edtech-awards.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Finalists and winners for &lt;a href=&quot;https://www.edtechdigest.com&quot;&gt;The EdTech Awards 2022&lt;/a&gt; have been announced to a worldwide audience of educators, technologists, students, parents, and policymakers interested in building a better future for learners and leaders in the education and workforce sectors. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The EdTech Awards were established in 2010 to recognise, acknowledge, and celebrate the most exceptional innovators, leaders, and trendsetters in education technology. Celebrating its 12th year, the US-based program is the world&apos;s largest recognition program for education technology, recognising the biggest names in edtech – and those who soon will be.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.edtechdigest.com/2022-finalists-winners/&quot;&gt;This year’s finalists and winners&lt;/a&gt; were narrowed from the larger field and judged based on various criteria, including: pedagogical workability, efficacy and results, support, clarity, value and potential.  &lt;/p&gt;
&lt;h4 &gt;I am absolutely delighted to announce that our work has been recognised in two categories this year:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;Pipeline Academy as finalist in the category Product or Service Setting a Trend&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;Daniel Molnar aka The Data Janitor as finalist in the category Educator Setting a Trend&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/the_edtech_awards-bltrfin22.png&quot; alt=&quot;The EdTech Awards Trendsetter Finalist 2022 badge&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;&quot;As events unfold on the world stage that seem to inch ever closer to a precipice unknown, we are reminded that the leaders and innovators of education technology have always worked on the edge,&quot; said Victor Rivero, who as Editor-in-Chief of EdTech Digest, oversees the program. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&quot;The future-focused work they do is inspired by the infinite potential of all people to learn and thrive. It&apos;s pushed forward by the human spirit. It&apos;s the light that even through the darkest times always shines through,&quot; Rivero said.&lt;/p&gt;
&lt;h4 &gt;On behalf of everyone who has been supporting the work of Pipeline Academy, we’d like to  say thank you for this recognition. It is proof that our efforts for making the education landscape a more transparent and future-proof place for learners does not go unnoticed by the broader innovation ecosystem.&lt;/h4&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - February 2022</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2022/03/02/the-data-janitor-letters-february-2022.html"/>
   <updated>2022-03-02T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2022/03/02/the-data-janitor-letters-february-2022.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://blog.fal.ai/the-unbundling-of-airflow-2/&quot; target=&quot;_blank&quot;&gt;The Unbundling of Airflow&lt;/a&gt; &lt;br&gt;&lt;em&gt;Gorkem Yurtseven, Co-Founder, Features and Labels&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A diverse set of tools is unbundling Airflow and this diversity is causing substantial fragmentation in modern data stack.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://dagster.io/blog/rebundling-the-data-platform&quot; target=&quot;_blank&quot;&gt;Rebundling the Data Platform&lt;/a&gt; &lt;br&gt;&lt;em&gt;Nick Schrock, Founder, Elementl&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A fundamentally new approach to orchestration that orients around assets rather than tasks.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://adat.blog/2022/02/fundraising-by-data-companies-in-2021/&quot; target=&quot;_blank&quot;&gt;Fundraising by data companies in 2021&lt;/a&gt; &lt;br&gt;&lt;em&gt;Bence Arató, Managing Director, BI Consulting&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;2021 was quite an exciting year in terms of funding for both data startups and established companies. We tracked more than a hundred data-related funding events during the year.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/kreuzwerker-gmbh/analytics-stacks-for-startups-ca5c131e3b29&quot; target=&quot;_blank&quot;&gt;Analytics Stacks for Startups&lt;/a&gt; &lt;br&gt;&lt;em&gt;Jan Katins, Senior IT Consultant/Data Engineer, kreuzwerker GmbH&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The stack should be relatively fast to implement (two weeks is possible), so you can quickly reap the benefits of having a data warehouse and BI Tooling in place or upload enriched data back to operational systems.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.datasciencesouth.com/blog/make/&quot; target=&quot;_blank&quot;&gt;Make Your Workflows Better With This Classic Unix Tool&lt;/a&gt; &lt;br&gt;&lt;em&gt;Adam Green, Data Engineer, Gridcognition&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Make and Makefiles for data pipelines.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.iccl.ie/news/gdpr-enforcer-rules-that-iab-europes-consent-popups-are-unlawful/&quot; target=&quot;_blank&quot;&gt;GDPR enforcer rules that IAB Europe&apos;s consent popups are unlawful&lt;/a&gt; &lt;br&gt;&lt;em&gt;Irish Council for Civil Liberties&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Google, Amazon, and the entire tracking industry relies on IAB Europe’s consent system, which has now been found to be illegal.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://noyb.eu/en/update-cnil-decides-eu-us-data-transfer-google-analytics-illegal&quot; target=&quot;_blank&quot;&gt;CNIL decides EU-US data transfer to Google Analytics illegal&lt;/a&gt; &lt;br&gt;&lt;em&gt;NOYB – European Center for Digital Rights&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Austrian and French Data Protection Authority: the continuous use of Google Analytics violates the GDPR.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://gizmodo.com/gdpr-iab-europe-privacy-consent-ad-tech-online-advertis-1848469604&quot; target=&quot;_blank&quot;&gt;Here&apos;s What&apos;s Wrong With GDPR&lt;/a&gt; &lt;br&gt;&lt;em&gt;Shoshana Wodinsky, Data Reporter, Gizmodo&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The EU&apos;s landmark privacy law, GDPR, was supposed to change the world of tech privacy forever. What the hell happened?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://spectrum.ieee.org/andrew-ng-data-centric-ai&quot; target=&quot;_blank&quot;&gt;Unbiggen AI&lt;/a&gt;&lt;br&gt;&lt;em&gt;Andrew Ng, Founder and CEO, Landing AI&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&quot;It’s time for smart-sized, “data-centric” solutions to big issues.&quot;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - January 2022</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2022/02/22/the-data-janitor-letters-january-2022.html"/>
   <updated>2022-02-22T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2022/02/22/the-data-janitor-letters-january-2022.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://mikkeldengsoe.substack.com/p/future-of-the-data-warehouse&quot; target=&quot;_blank&quot;&gt;We’ve only scratched the surface of the full potential for the data warehouse&lt;/a&gt; &lt;br&gt;&lt;em&gt;Mikkel Dengsøe, Head of Data Science, Operations &amp;amp; Financial Crime, Monzo Bank&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Why I think the data warehouse will become the control centre for modern companies&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://vickiboykis.com/2022/01/09/git-sql-cli/&quot; target=&quot;_blank&quot;&gt;Git, SQL, CLI&lt;/a&gt; &lt;br&gt;&lt;em&gt;Vicki Boykis, Machine Learning Engineer, Automattic&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I’ve narrowed it down to three basic tools.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
One Year of dbt &lt;br&gt;&lt;em&gt;Adam Boscarino, Director of Data Engineering, Devoted Health&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Over the last year, dbt has become a key piece of the data platform at Devoted and lived up to our wildest hopes and dreams. We have gone from a single proof-of-concept to 1,100+ models.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.open-metadata.org/why-openmetadata-is-the-right-choice-for-you-59e329163cac&quot; target=&quot;_blank&quot;&gt;Why OpenMetadata is the Right Choice for you&lt;/a&gt; &lt;br&gt;&lt;em&gt;Suresh Srinivas, Co-Founder, Collate&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Is OpenMetadata pull-based, push-based, or hybrid? Again, all systems are mainly pull-based to integrate with metadata sources. We support push-based ingestion when it is possible.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://groupby1.substack.com/p/data-engineering&quot; target=&quot;_blank&quot;&gt;The future history of Data Engineering&lt;/a&gt; &lt;br&gt;&lt;em&gt;Matt Arderne, Product Engineering, Data &amp;amp; Analytics, focal&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;On Data Engineers and their place in a Data SaaS world.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://hightouch.io/blog/airflow-alternatives-a-look-at-prefect-and-dagster/&quot; target=&quot;_blank&quot;&gt;Airflow Alternatives: A Look at Prefect and Dagster&lt;/a&gt; &lt;br&gt;&lt;em&gt;Pedram Navid, Head of Data, Hightouch&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We take a deep dive into Airflow, Prefect, and Dagster and the differences between the three!&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://towardsdatascience.com/modern-data-stack-which-place-for-spark-8e10365a8772&quot; target=&quot;_blank&quot;&gt;Modern Data Stack: Which Place for Spark?&lt;/a&gt; &lt;br&gt;&lt;em&gt;Furcy Pin, Lead Data Engineer, Younited Credit&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This makes data-lineage more difficult, since dbt only lets us visualize the BigQuery part, while our internal “dbt for pySpark” tool only lets us see the pySpark part.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20220313161240/https://nirantk.com/writing/why-i-quit-data-science.html&quot; target=&quot;_blank&quot;&gt;Why I Quit Data Science&lt;/a&gt; &lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Nirant Kasliwal, Analytics Engineer, Sundial&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Question from a friend: I am interested in knowing how did you come to this decision of moving to SWE from DS/MLE.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Become A Better Data Engineer On A Shoestring (More Free Resources)</title>
   <link href="https://dataengineering.academy/2022/02/18/learn-data-engineering-on-a-shoestring-free-courses.html"/>
   <updated>2022-02-18T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2022/02/18/learn-data-engineering-on-a-shoestring-free-courses.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;A bit more than a year ago I’ve compiled an annotated list of &lt;a href=&quot;/2020/12/15/become-a-data-engineer-on-a-shoestring.html&quot;&gt;the best free courses and learning resources&lt;/a&gt; that could help anyone to become a data engineer on a shoestring. We’ve received an overwhelming amount of positive feedback on it, so after a full year of running the bootcamp I sat down again and collected an other bunch of resources we’ve bumped into during the cohorts.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Experienced data scientists, ambitious data analysts, data-obsessed product managers and future-oriented computer scientist have joined Pipeline Academy in the last year with the goal of learning real data craftsmanship, aka data engineering. But even if you don’t identify as any of the above, but you’re planning to become a more well-rounded data professional, you’ve come to the right place.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Expect well-designed learning experiences that are mostly very accessible in terms of pricing, and learning outcomes that are very close to what the market demands.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Some guidance for this list:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;If you haven’t read &lt;a href=&quot;/2020/12/15/become-a-data-engineer-on-a-shoestring.html&quot;&gt;the first edition of our compiled list of learning materials&lt;/a&gt;, make sure to check it out first.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;There are various different formats included here (video-based courses, books, podcasts, story-based interactive coding tutorials etc.), try to identify which ones suit your learning style the best.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Don’t skip the fundamentals: using Python, SQL and the command line are essential for data engineers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Integrate learning into your weekly schedule: try sticking to what you’ve started.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Share your experience and recommendations with others, it’s really difficult to find the right courses in the forest of mediocre Medium posts and useless certifications.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;General&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.dataengineeringpodcast.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;The Data Engineering Podcast&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; and &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://www.pythonpodcast.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;The Python Podcast.init&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;Outstanding podcasts of Tobias Macey on the mentioned topics.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://web.archive.org/web/20220301165655/https://cal-data-eng.github.io/&quot;&gt;&lt;strong&gt;&lt;em&gt;Data Engineering, UC Berkeley, Spring 2021&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;University course ran by industry players with experience, although one has to reflect on the fact that data engineering does not happen in notebooks.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://technically.substack.com/archive&quot;&gt;&lt;strong&gt;&lt;em&gt;Technically&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Technically sends out engaging, simple explanations of technical concepts that are useful for your day to day job and fun to read. Start with the &lt;a href=&quot;https://technically.dev/posts/sql-for-the-rest-of-us&quot;&gt;SQL&lt;/a&gt; and the &lt;a href=&quot;https://technically.dev/posts/aws-for-the-rest-of-us&quot;&gt;AWS&lt;/a&gt; explanations.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;http://highscalability.com/blog/2017/12/18/explain-the-cloud-like-im-10.html&quot;&gt;&lt;strong&gt;&lt;em&gt;Explain the Cloud Like I&apos;m 10&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Beginners will find the cloud explained from the basics. Little prior knowledge is assumed. You will find lots of pictures, lots of examples, and many somewhat questionable analogies in this book.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;&lt;em&gt;Alex Xu: System Design Interview - An Insider&apos;s Guide&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;The &lt;a href=&quot;https://www.amazon.de/-/en/Alex-Xu/dp/B08CMF2CQF&quot;&gt;book&lt;/a&gt; and the &lt;a href=&quot;https://courses.systeminterview.com/courses/system-design-interview-an-insider-s-guide&quot;&gt;course&lt;/a&gt; will help you to think about and understand complex integrations.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/unsplash-image-y7d265_7i08.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Python&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://allendowney.github.io/DSIRP/&quot;&gt;&lt;strong&gt;&lt;em&gt;Data Structures and Information Retrieval in Python&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;The new Downey book introduces data structures and algorithms using a web search engine as a motivating example.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://realpython.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;Real Python&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;The best Python tutorials — even the free ones are outstanding.&lt;/p&gt;
&lt;h4 &gt;SQL&lt;/h4&gt;
&lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Update: our advisor, &lt;a href=&quot;https://martin-loetzsch.de&quot;&gt;Dr. Martin Loetzsch&lt;/a&gt; just shared for free the material he&apos;s teaching also at Pipeline Data Engineering Academy.&lt;/p&gt;
&lt;/div&gt;
&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/8HlNG8bdlM0?si=oKya_Qzs6WU-826o&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/24Uvo5vZJWA?si=jVvcl4xSqQGB8oM4&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://opentextbc.ca/dbdesign01/&quot;&gt;&lt;strong&gt;&lt;em&gt;Database Design&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;2nd Edition from The BC Open Textbook Project.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://github.com/oleg-agapov/data-engineering-book&quot;&gt;&lt;strong&gt;&lt;em&gt;oleg-agapov/data-engineering-book&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;A good, partial start for SQL from Oleg Agapov in the &lt;code&gt;Beginner path&lt;/code&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.youtube.com/channel/UCHnBsf2rH-K7pn09rb3qvkA&quot;&gt;&lt;strong&gt;&lt;em&gt;CMU Database Group&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;Video series on contemporary databases and datawarehouses.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/learn_data_engineering_course_pipeline_academy.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;CLI&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The command line can be intimidating if you&apos;re starting out with it. Two resources that are nice to novices especially if they don&apos;t come from a computer science background.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://programminghistorian.org/&quot;&gt;&lt;strong&gt;&lt;em&gt;Programming Historian&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Example: &lt;a href=&quot;https://programminghistorian.org/en/lessons/json-and-jq&quot;&gt;Reshaping JSON with jq&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://software-carpentry.org/&quot;&gt;&lt;strong&gt;&lt;em&gt;Software Carpentry&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Example: &lt;a href=&quot;https://swcarpentry.github.io/make-novice/&quot;&gt;Automation and Make&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://github.com/veltman/clmystery&quot;&gt;&lt;strong&gt;&lt;em&gt;veltman/clmystery&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;The Command Line Murders - an interactive learnbook.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://gitlab.com/slackermedia/bashcrawl&quot;&gt;&lt;strong&gt;&lt;em&gt;slackermedia / bashcrawl&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;Learn Linux commands by playing a simple text adventure.&lt;/p&gt;
&lt;h4 &gt;DevOps&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://acloudguru.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;A Cloud Guru&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;Good choice for paid courses on the cloud and DevOps.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.gentlydownthe.stream/&quot;&gt;&lt;strong&gt;&lt;em&gt;Gently Down the Stream&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;A gentle introduction to Apache Kafka - pairing whimsical imagery with lucid explanations of stream processing concepts, this book will captivate beginners of all ages.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.youtube.com/watch?v=4ht22ReBjno&quot;&gt;&lt;strong&gt;&lt;em&gt;The Illustrated Children&apos;s Guide to Kubernetes&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;br&gt;Follow the adventures of Phippy the Giraffe, Captain Kube, and Goldie the Gopher as they discover Kubernetes pods, replication controllers, services, and volumes. Get silly and serious at the same time with this lighthearted introduction to core Kubernetes concepts.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Thanks to all graduates and expert guests of our cohorts in 2021 for their valuable feedback.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Happy learning!&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>What Did You Build at Pipeline Academy? This.</title>
   <link href="https://dataengineering.academy/2022/02/18/data-engineering-capstone-project-data-product.html"/>
   <updated>2022-02-18T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2022/02/18/data-engineering-capstone-project-data-product.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineers have to wear many different hats at the same time: they are architects, designers, builders, maintainers, procurement and quality assurance — to just name a few. If you’d like to break into this profession, you need to prove that you can do all of the above, and more.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;One of the key assets you can use to do that is a data product that you’ve built with your own hands. Last summer we’ve already covered the expectations towards a portfolio project in detail, focusing on a general outline and approach. Take a good look: &lt;/p&gt;
&lt;blockquote&gt;&lt;h4 &gt;Recommended READ:&lt;br&gt;&lt;a href=&quot;/2021/07/02/the-data-engineering-portfolio-project.html&quot;&gt;The data engineering Portfolio project&lt;/a&gt;
&lt;/h4&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;If you’ve never built a data stack from scratch, it’s difficult to imagine how it’s done. Especially when you face a timeline of 12 days total without any real buffers.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/unsplash-image-wx2l8l-fgea.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Meanwhile, we see the following patterns emerge:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;some participants come with a clear project idea, they would like to put a machine learning model into production, or would like to revamp the supporting infrastructure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;some come with a data stack problem that they are currently facing at work, and they need professional guidance and clarity about how to go about solving it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;furthermore - and this might be the most frequent setup - some people come without any clue about what they really wan’t to build. And that’s totally fine.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Below you can find seven outstanding examples for the capstone projects of our alumni. They include air quality monitoring IoT solutions, maps for tracking invasive species in Europe, a Billboard chart for the hottest NFTs and so much more… &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;They are a great showcase of what you can do with the technologies you learn in our course. Without any further ado, here’s a short list of the projects and products our graduates have built during 2021.&lt;/p&gt;
&lt;h4 &gt;Portfolio projects&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Kiril Kasjanov: &lt;a href=&quot;https://github.com/kirajano/wiki-streams&quot;&gt;Wikipedia Event Streaming and Real-Time Analytics with Kafka + ElasticStack &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Patrycja Ottawa: &lt;a href=&quot;https://github.com/PatrycjaO/PipelineProject&quot;&gt;Wind turbine monitoring&lt;/a&gt; &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Sarah Ni: NFTcharts: &lt;a href=&quot;https://github.com/sni-c/final-project&quot;&gt;Uncovering Trade Activity &amp;amp; Hot Trends &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Widad Iqbal Mogral: &lt;a href=&quot;https://github.com/widadmogral/Invasive-species-in-EU-Mapper&quot;&gt;Invasive Species in EU Mapper &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Lukas Hondrich: &lt;a href=&quot;https://github.com/lukashondrich/twitterradio&quot;&gt;Realtime-ish Twitter Radio &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Florian Motz: &lt;a href=&quot;https://github.com/FloMoTionGo/Capstone-Project-PA2021&quot;&gt;eMobility availability at VBB Points Of Interests &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Romina Nikolova: &lt;a href=&quot;https://github.com/romina-nikolova/free-air-quality&quot;&gt;Free Air Quality Monitoring &lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Interested in building something similar with the guidance of industry experts? Just hit the &lt;a href=&quot;/apply/&quot;&gt;apply button&lt;/a&gt; and let’s have a chat!&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - December 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2022/01/18/the-data-janitor-letters-december-2021.html"/>
   <updated>2022-01-18T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2022/01/18/the-data-janitor-letters-december-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://ottertune.com/blog/2021-databases-retrospective/&quot; target=&quot;_blank&quot;&gt;Databases in 2021: A Year in Review&lt;/a&gt; &lt;br&gt;&lt;em&gt;Dr. Andy Pavlo, Co-founder, OtterTune&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It was a wild year for the database industry, with newcomers overtaking the old guard, vendors fighting over benchmark numbers, and eye-popping funding rounds. We also had to say goodbye to some of our database friends through acquisitions, bankruptcies, or retractions.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20221215182441/https://towardsdatascience.com/building-an-end-to-end-open-source-modern-data-platform-c906be2f31bd&quot; target=&quot;_blank&quot;&gt;Building an End-to-End Open-Source Modern Data Platform&lt;/a&gt; &lt;br&gt;&lt;em&gt;Mahdi Karabiben, Senior Data Engineer, Zendesk&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A detailed guide to help you navigate the modern data stack and build your own platform using open-source technologies.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.mattritter.me/?p=398&quot; target=&quot;_blank&quot;&gt;Postgres, Kafka, and the Mysterious 100 GB&lt;/a&gt; &lt;br&gt;&lt;em&gt;Matt Ritter, Software Engineer, Google&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I think the experience shows that PoCs using “production infrastructure” can expose pitfalls that might appear in a real implementation.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.wesleyac.com/posts/consider-sqlite&quot; target=&quot;_blank&quot;&gt;Consider SQLite&lt;/a&gt; &lt;br&gt;&lt;em&gt;Wesley Aptekar-Cassels&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;As long as you don&apos;t expect to need tens of thousands of small writes per second, thousands of large writes, or long-lived write transactions, it&apos;s highly likely that SQLite will support your usecase.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.crunchydata.com/blog/five-tips-for-a-healthier-postgres-database-in-the-new-year&quot; target=&quot;_blank&quot;&gt;Five Tips For a Healthier Postgres Database in the New Year&lt;/a&gt; &lt;br&gt;&lt;em&gt;Craig Kerstiens, Product, Crunchy Data&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;While onboarding customer after customer this year I&apos;ve noted a few key things everyone should put in place right away - to either improve the health of your database or to save yourself from a bad day.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://petewarden.com/2021/12/11/the-death-of-feature-engineering-is-greatly-exaggerated/&quot; target=&quot;_blank&quot;&gt;The Death of Feature Engineering is Greatly Exaggerated&lt;/a&gt; &lt;br&gt;&lt;em&gt;Pete Warden, Staff Research Engineer, Google&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I don’t want to minimize deep learning’s achievements in reducing the toil involved in building feature pipelines, I’m still constantly amazed at how effective they are. I would like to see more emphasis put on feature engineering in research and teaching though.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://arnoldgalovics.com/microservices-in-production/&quot; target=&quot;_blank&quot;&gt;Don’t start with microservices in production – monoliths are your friend&lt;/a&gt; &lt;em&gt;&lt;br&gt;Arnold Gálovics, Delivery Manager, EPAM Systems &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I’m gonna hit you with hard truth my friend.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://christine.website/blog/open-source-broken-2021-12-11&quot; target=&quot;_blank&quot;&gt;&quot;Open Source&quot; is Broken&lt;/a&gt; &lt;br&gt;&lt;em&gt;Christine Dodrill, Software Designer, Tailscale&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Why I don&apos;t write useful software unless you pay me.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Our 2021 in a Nutshell</title>
   <link href="https://dataengineering.academy/2021/12/20/2021-data-engineering-in-a-nutshell-2022.html"/>
   <updated>2021-12-20T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/12/20/2021-data-engineering-in-a-nutshell-2022.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;It&apos;s the time of the year when everybody is trying to summarise what happened in the last 12 months: &apos;best of&apos; lists, highlights of the year and predictions for 2022 are dominating your inbox. This blog post is not different either.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;/2021/01/28/2020-a-year-to-remember.html&quot;&gt;2020 was definitely eventful&lt;/a&gt;, and 2021 came with its own set of surprises. But Pipeline Academy finally managed to get off the ground, we&apos;ve launched three amazing cohorts and had loads of fun together with people from across the globe — literally. &lt;a href=&quot;/2021/06/29/pipeline-academy-on-the-data-engineering-podcast-vol2.html&quot;&gt;The chat with Tobias Macey on the Data Engineering Podcast&lt;/a&gt; serves as memento for how we&apos;ve felt after our first cohort had graduated.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Other than that, here are just a few of the highlights that shaped our year:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;We&apos;ve spent the year at our temporary home, &lt;a href=&quot;/2021/03/10/the-campus.html&quot;&gt;the CIEE campus&lt;/a&gt;, but as you are reading these words we&apos;re already moving in to our new office (more on this soon).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;More and more organisations realise how difficult it is to recruit people with data engineering competencies, or approach us to support them &lt;a href=&quot;/2021/06/08/data-sustainability-zsofia-bognar-from-ecosia.html&quot;&gt;on their way of building compliant and sustainable data infrastructures&lt;/a&gt;: if you can identify with any of this, feel free to message us.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The &quot;&lt;em&gt;here is how it&apos;s done, repeat after me&lt;/em&gt;&quot; type of education a lot of data courses provide does not work for data engineering. Read some insights on &lt;a href=&quot;/2021/07/26/how-to-teach-data-engineering.html&quot;&gt;how we think about education&lt;/a&gt; and &lt;a href=&quot;/2021/07/02/the-data-engineering-portfolio-project.html&quot;&gt;building data products&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;While &lt;a href=&quot;/2021/10/08/meet-the-graduates-michele-tassoni.html&quot;&gt;our first graduates&lt;/a&gt; are already out and about, &lt;a href=&quot;/apply/&quot;&gt;we&apos;re busy launching our part-time course in a US-friendly timezone&lt;/a&gt;!&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The winners of this year’s Pipeline Academy Awards (aka The Pipies) have been announced. What the hell is this about? &lt;a href=&quot;/2021/12/14/the-pipeline-academy-awards-2021-pipies.html&quot;&gt;Read here&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;But at the end of the day, &lt;a href=&quot;https://www.coursereport.com/schools/pipeline-data-engineering-academy&quot;&gt;the kind words from our graduates&lt;/a&gt; are what makes us proud.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Data engineering as an economic force and a career path is skyrocketing, thousands of people in tech and data realise how crucial these technical skills are. If you&apos;re struggling with half-baked online courses, or if you&apos;re looking for a significant increase in your salary... you know where to find us.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Thanks to all supporters, expert speakers, clients and participants for having our backs.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Happy holidays and a happy new year. Stay safe and healthy.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Best,&lt;br&gt;Daniel and Peter&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/unsplash-image-sutffcahv_a.jpg&quot; alt=&quot;&quot;&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - November 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/12/15/the-data-janitor-letters-november-2021.html"/>
   <updated>2021-12-15T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/12/15/the-data-janitor-letters-november-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.iccl.ie/news/online-consent-pop-ups-used-by-google-and-other-tech-firms-declared-illegal/&quot; target=&quot;_blank&quot;&gt;Tracking-industry body IAB Europe told that it has infringed the GDPR&lt;/a&gt;&lt;br&gt;&lt;em&gt;Irish Council for Civil Liberties&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Google and the entire tracking industry relies on IAB Europe’s consent system, which has now been found to be illegal.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://petrjanda.substack.com/p/why-the-data-analyst-role-has-never&quot; target=&quot;_blank&quot;&gt;Why the Data Analyst role has never been harder&lt;/a&gt;&lt;br&gt;&lt;em&gt;Petr Janda, CTO, Pleo&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The curse of complexity around the Modern Data Stack.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://mikkeldengsoe.substack.com/p/spreadsheets-in-the-data-stack&quot; target=&quot;_blank&quot;&gt;Spreadsheets: The Duct Tape of the Modern Data Stack&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mikkel Dengsøe, Head of Data Science, Operations &amp;amp; Financial Crime, Monzo Bank&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Spreadsheets are the interface that allows anyone to quickly and easily bring data into the data warehouse.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://robertsahlin.com/serverless-dbt-on-google-cloud-platform/&quot; target=&quot;_blank&quot;&gt;Serverless dbt on Google Cloud Platform&lt;/a&gt;&lt;br&gt;&lt;em&gt;Robert Sahlin, Senior Data Engineer, &lt;/em&gt;&lt;a href=&quot;http://mathem.se/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;MatHem.se&lt;/em&gt;&lt;/a&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A serverless solution to run dbt in a self-hosted and collaborative setup and being able to follow GitOps style.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/the-prefect-blog/orchestrating-elt-with-prefect-and-dbt-a-flow-of-flows-part-1-aac77126473&quot; target=&quot;_blank&quot;&gt;Orchestrating ELT with Prefect and dbt — a Flow of Flows (Part 1)&lt;/a&gt; &lt;br&gt;&lt;em&gt;Anna Geller, Solutions Engineer, Prefect&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;How to manage dependencies between data pipelines.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://servian.dev/modelling-type-1-2-slowly-changing-dimensions-with-dbt-1b80078f290a&quot; target=&quot;_blank&quot;&gt;Modelling Type 1 + 2 Slowly Changing Dimensions with dbt&lt;/a&gt; &lt;br&gt;&lt;em&gt;Weng-Kin Lee, Associate Consultant, Servian&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;By following this pattern to create dimension models, it is easy to incorporate both Type 1 and Type 2 changes into the same dimension. Using dbt macros, we have modularised the dimension’s functionality for reusability and provide capacity to add more functionality to the dimension.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/data-monzo/mapping-our-data-journey-with-column-lineage-56209c00606d&quot; target=&quot;_blank&quot;&gt;Mapping our data journey with column lineage&lt;/a&gt; &lt;br&gt;&lt;em&gt;Borja Vázquez Barreiros, Senior Analytics Engineer, Monzo&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;At the time of writing, we have over 4700 data models in our production dbt project, and over 800 views defined in Looker 🤯.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://hakibenita.com/postgresql-unknown-features&quot; target=&quot;_blank&quot;&gt;Lesser Known PostgreSQL Features&lt;/a&gt; &lt;br&gt;&lt;em&gt;Haki Benita, Development Team Lead, PCENTRA&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Most of us are not aware of all the features in tools we use on a daily basis, especially if it&apos;s big and extensive like PostgreSQL.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://duckdb.org/2021/10/29/duckdb-wasm.html&quot; target=&quot;_blank&quot;&gt;DuckDB-Wasm: Efficient Analytical SQL in the Browser&lt;/a&gt; &lt;br&gt;&lt;em&gt;André Kohn and Dominik Moritz, DuckDB&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It is powered by WebAssembly, speaks Arrow fluently, reads Parquet, CSV and JSON files backed by Filesystem APIs or HTTP requests and has been tested with Chrome, Firefox, Safari and Node.js.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://thundergolfer.com/kubernetes/infrastructure/data-engineering/2021/11/04/from-data-eng-to-sys-admin-put-down-k8s/&quot; target=&quot;_blank&quot;&gt;From Data Engineer to SysAdmin&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jonathon Belotti, Team Lead Data Platform, Canva&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Put down the K8s cluster, your pipelines can run without it.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.replit.com/nix-vs-docker&quot; target=&quot;_blank&quot;&gt;Will Nix Overtake Docker?&lt;/a&gt; &lt;br&gt;&lt;em&gt;Connor Brewster, Software Engineer, Replit&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;No, these tools accomplish different goals, however they can be used in combination to provide the best of both worlds: reproducible builds and containerized deployments.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://erikbern.com/2021/11/30/storm-in-the-stratosphere-how-the-cloud-will-be-reshuffled.html&quot; target=&quot;_blank&quot;&gt;Storm in the stratosphere: how the cloud will be reshuffled&lt;/a&gt; &lt;br&gt;&lt;em&gt;Erik Bernhardsson&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Cloud vendors will increasingly focus on the lowest layers in the stack: basically leasing capacity in their data centers through an API. Other pure-software providers will build all the stuff on top of it. Databases, running code, you name it.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Pipeline Academy Awards 2021</title>
   <link href="https://dataengineering.academy/2021/12/14/the-pipeline-academy-awards-2021-pipies.html"/>
   <updated>2021-12-14T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/12/14/the-pipeline-academy-awards-2021-pipies.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;The end of the year has arrived, our first full year in operation. 2021 was about the Cambrian explosion of data engineering tooling, yet you don&apos;t have to be a data scientist to be certain that 90% of the data tools will be gone in about two years or so, and for a good reason. Like it or not, most of them are &lt;a href=&quot;https://twitter.com/soobrosa/status/1463509394969223175&quot;&gt;solutions for fictional problems&lt;/a&gt;. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Whether you are a seasoned data architect or a newcomer in the world of data infrastructure, you already know that a major part of your work is about curating the right tools for your business, your team and your data. One of the key assets that our participants leave the course with is a framework for making informed decisions about data infrastructure tooling. It considers timeless software engineering practices, state-of-the-art-technologies, business goals and even sustainability factors, but under the hood it is very much powered by common sense.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This framework (and the desire to avoid major migration projects) is what prevents data engineers from being misled by the shockwave of marketing messages that aim to convince them that a new and shiny tool is a must have in... wait for it...&lt;/p&gt;
&lt;blockquote&gt;&lt;h2 &gt;THE MODERN DATA STACK.&lt;/h2&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;We deal with the process of making these tooling decisions on a daily basis. We cover a lot of them in the course, teach them, love them, update them and use them. And we have favourites that stand out.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The 2021 Pipeline Academy Awards (The Pipies) are brought to you by Pipeline Academy. It&apos;s our way of praising the teams that built meaningful software solutions that serve a real purpose, the ones that are most likely here to stay to make the lives of data engineers better.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/awards_pipies_pipeline_academy.png&quot; alt=&quot;The Pipies 2021 award badge&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Separating the wheat from the chaff&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Generally what we value:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;the tool is open source, you can run it on your own if you have the resources and the expertise,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;the tool has a paid hosted option for the occasion that you don&apos;t have the means to host it on your own,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;the hosted option has a generous free tier to start with,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;good tutorials and documentation is provided,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;learning curve is not steep.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Caveat: the Apple M1 hell is real, the support for it is not really there yet. &lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h1 &gt;The 2021 Pipies go to…&lt;/h1&gt;
&lt;h4 &gt;Data Acquisition: &lt;a href=&quot;https://airbyte.io/&quot;&gt;Airbyte&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Open-Source Data Integration Pipelines | ELT&lt;/code&gt;. Our choice for EL, plays nice with &lt;code&gt;dbt&lt;/code&gt;, bit heavy on the resource side, Docker as a Lambda :)&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/airbyte_pipeline_academy.jpg&quot; alt=&quot;Airbyte, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Telemetry: &lt;a href=&quot;https://quix.ai/&quot;&gt;Quix&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Real-time stream processing PaaS&lt;/code&gt;. Whatever takes the pain out of Kafka is a friend of ours. A European player with McLaren expertise, supported by &lt;a href=&quot;https://www.project-a.com/&quot;&gt;Project A&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/quix_pipeline_academy.png&quot; alt=&quot;Quix, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;ETL/ELT: &lt;a href=&quot;https://www.prefect.io/&quot;&gt;Prefect&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;The New Standard in Dataflow Automation&lt;/code&gt;. &lt;a href=&quot;https://www.reddit.com/r/ExperiencedDevs/comments/nmodyl/drunk_post_things_ive_learned_as_a_sr_engineer/&quot;&gt;&quot;Airflow is ****, yes.&quot;&lt;/a&gt;, so &lt;a href=&quot;https://medium.com/the-prefect-blog/why-not-airflow-4cfa423299c4&quot;&gt;Why Not Airflow?&lt;/a&gt;. A product that evolves in a great way, already at its second-generation workflow engine, &lt;a href=&quot;https://www.prefect.io/blog/announcing-prefect-orion&quot;&gt;Orion&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/prefect_pipeline_academy_pipies.png&quot; alt=&quot;Prefect, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Data Modeling: &lt;a href=&quot;https://www.getdbt.com/&quot;&gt;dbt&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Transform data in your warehouse&lt;/code&gt;. The whale that rules the T in ELT that coined the term &lt;a href=&quot;https://www.getdbt.com/what-is-analytics-engineering/&quot;&gt;Analytics Engineering&lt;/a&gt;. dbt encodes best practices, a product built on experience solving actual problems; the jury is still out there on how it scales for BIG data though. Still should be good enough for 90% of the BI teams.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/screenshot-2021-12-10-at-14.24.49.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Data Warehousing: &lt;a href=&quot;https://clickhouse.com/&quot;&gt;ClickHouse&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;fast open-source OLAP DBMS&lt;/code&gt;. From Russia with love, written in C, just &lt;a href=&quot;https://tech.marksblogg.com/benchmarks.html&quot;&gt;a magnitude faster&lt;/a&gt; than anything else especially if it starts with Apache.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/clickhouse_pipeline_academy_pipies.jpg&quot; alt=&quot;ClickHouse, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Data Quality: &lt;a href=&quot;https://greatexpectations.io/&quot;&gt;Great Expectations&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;a shared, open standard for data quality&lt;/code&gt;. The Janus-faced power tool, marrying the best of both machine-readibility and human oversight. Outstanding learning curve and workflow, can get heavy on deployment because of its massive dependencies.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/great_expectations_pipeline_academy.jpg&quot; alt=&quot;great_expectations, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Clouds: &lt;a href=&quot;https://www.pulumi.com/&quot;&gt;Pulumi&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Modern Infrastructure as Code&lt;/code&gt;. Infrastructure as code, not configfiles, no YAML horror - even in Python. Period.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/pulumi_pipeline_academy.jpg&quot; alt=&quot;Pulumi, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Deployment: &lt;a href=&quot;https://fly.io/&quot;&gt;Fly&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Deploy app servers close to your users&lt;/code&gt;. &lt;a href=&quot;https://fly.io/blog/docker-without-docker/&quot;&gt;Dark magic&lt;/a&gt; to deploy your Docker images in a whim.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/screenshot-2021-12-10-at-14.54.53.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Serving Data: &lt;a href=&quot;https://streamlit.io/&quot;&gt;Streamlit&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;The fastest way to build and share data apps&lt;/code&gt;. Shiny for Python, straps a beautiful interactive interface on your Python logic.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/streamlit_pipeline_academy.jpg&quot; alt=&quot;Streamlit, a Pipies 2021 winner&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Honorable mentions:&lt;/h3&gt;
&lt;h4 &gt;Computation: &lt;a href=&quot;https://saturncloud.io/&quot;&gt;Saturn Cloud&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Data Science &amp;amp; Machine Learning with Dask &amp;amp; GPUs&lt;/code&gt;. I still like the old tagline: &lt;code&gt;It&apos;s like Spark. Except you won&apos;t hate yourself.&lt;/code&gt;&lt;/p&gt;
&lt;h4 &gt;CD/CI: &lt;a href=&quot;https://github.com/features/actions&quot;&gt;GitHub Actions&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;Automate your workflow from idea to production&lt;/code&gt;. The only thing that I feel “&lt;code&gt;how-could-we-live-without-it” about&lt;/code&gt;, also &lt;a href=&quot;https://simonwillison.net/2021/Mar/5/git-scraping/&quot;&gt;git scraping&lt;/a&gt;.&lt;/p&gt;
&lt;h4 &gt;Serving Data: &lt;a href=&quot;https://datasette.io/&quot;&gt;Datasette&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Self-definition: &lt;code&gt;An open source multi-tool for exploring and publishing data&lt;/code&gt;. Mentioning Simon Willison, let&apos;s talk about the power of SQLite, an interface, all possibly Dockerized with a few lines of code.&lt;/p&gt;
&lt;h3 &gt;
&lt;br&gt;The Pipies 2021&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Big shout-out and thank you to all of the teams who are responsible for these products, it&apos;s a pleasure for us to have so many amazing possibilities to choose from.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;If you haven&apos;t yet, we recommend every data professional to check out these tools and compare them with the solutions you&apos;re already familiar with. Let us know what you think about the winners, and feel free to send us your preferred data tools and updates from this year!&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;See you in 2022, when The Pipies return.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;Disclosure: Pipeline Academy is not financially affiliated with any of the above organisations.&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Skills Gap in Data Engineering</title>
   <link href="https://dataengineering.academy/2021/11/24/data-engineering-skills-gap.html"/>
   <updated>2021-11-24T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/11/24/data-engineering-skills-gap.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Most data professionals realise very early in their journey that accessing the knowledge that they really need to solve data engineering problems is hard to come by. The other thing they don’t necessarily see is how short-sighted a lot of courses are, and how most of the technical content they provide is going to be rendered useless in a year or two.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Whenever people talk about the so called skills gap, they focus on addressing problems in a systemic scale and try to come up with a solution that is meant to address multiple dimensions at the same time: time, social factors, economic trends, technical advancement, legal changes just to name a few. This is complicated by the conflicting agendas of the involved stakeholders: employers, learners, government institutions and lastly, educational institutions.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/unsplash-image-e3tdq04ns2s.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;I was lucky enough to be involved in an action group organised by &lt;a href=&quot;https://emerge.education/&quot;&gt;&lt;span style=&quot;text-decoration:underline&quot;&gt;Emerge&lt;/span&gt;&lt;/a&gt;, &lt;a href=&quot;https://www.coursera.org/&quot;&gt;&lt;span style=&quot;text-decoration:underline&quot;&gt;Coursera&lt;/span&gt;&lt;/a&gt;, &lt;a href=&quot;https://ufi.co.uk&quot;&gt;Ufi&lt;/a&gt; and &lt;a href=&quot;https://learn.filtered.com/&quot;&gt;&lt;span style=&quot;text-decoration:underline&quot;&gt;Filtered&lt;/span&gt;&lt;/a&gt;, that set itself the goal of coming up with a &lt;a href=&quot;https://ufi.co.uk/latest/future-proofing-skills-development/&quot;&gt;green paper that “examines how we can develop the necessary skills and establish new pathways into jobs”&lt;/a&gt;. My contribution in this work was very humble, but  this should not stop you from reading the study. I don’t want to miss the opportunity to share this with you so everyone can have a sneak peek into the world of structural challenges that educators and edtech entrepreneurs have to tackle in the upcoming years.&lt;/p&gt;
&lt;h4 &gt;READ THE FULL REPORT: &lt;a href=&quot;https://ufi.co.uk/latest/future-proofing-skills-development/&quot;&gt;&lt;span style=&quot;text-decoration:underline&quot;&gt;Developing skills and establishing new pathways into jobs (.pdf)&lt;/span&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;You see, there is a strategic and a tactical level to addressing the challenges of the near future, but they can’t and shouldn’t be separated. Otherwise this will lead to fundamentally broken outcomes for learners: using wrong didactical methods, developing meaningless curricula, abusing technology for the wrong purposes, just to name a few. But false marketing messages (“&lt;em&gt;profession of the next decade&lt;/em&gt;” — does this ring a bell?) are just as much of a problem, they can lead masses towards an unsustainable career direction.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Long story short, ultimately I tend to come back to this quote from John Alexander Smith, Professor of Moral Philosophy at Oxford from 1914:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Gentlemen, you are now about to embark on a course of studies that (will) form a noble adventure…Let me make this clear to you. ..nothing that you will learn in the course of your studies will be of the slightest possible use to you in after life – save only this – that if you work hard and intelligently, you should be able to detect when a man is talking rot, and that, in my view, is the main, if not the sole purpose of education.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;The faster the world moves, the more you’ll have to learn to avoid being misled. Lifelong learning is not a trend.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>10 Reasons Why Aspiring Data Engineers Choose Pipeline Academy</title>
   <link href="https://dataengineering.academy/2021/11/09/10-reasons-learn-data-engineers-choose-pipeline-academy-course-training.html"/>
   <updated>2021-11-09T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/11/09/10-reasons-learn-data-engineers-choose-pipeline-academy-course-training.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Ambitious data analysts, data scientists trying to take their careers to the next level, product owners aiming to build next-generation data products, software engineers dealing with legacy data stacks... they all are facing the same challenge: how do I get the data engineering skills that enable me to achieve my goals?&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This frustration is very real, and it is indeed very common! Plenty of professionals are trying to find a simple and well-structured leaning path for data engineering, but the road is paved with deceptive marketing messages and fake experts.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The students who ended up at Pipeline Academy gave us their reasons for picking us, and this list was born. You can &lt;a href=&quot;https://www.coursereport.com/schools/pipeline-data-engineering-academy&quot;&gt;read their reviews on CourseReport&lt;/a&gt;. So if you feel that you can identify with any of the below, just &lt;a href=&quot;/apply/&quot;&gt;schedule a call and have a chat with me&lt;/a&gt;!&lt;/p&gt;
&lt;h4 &gt;1. Online courses don’t work for most&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Raise your hand if you&apos;ve finished every online MOOC course you&apos;ve started. Trust me, you are not alone. It&apos;s difficult to filter for the meaningful platforms/courses that deliver value especially when it comes to learning data engineering, but even if you find them (&lt;a href=&quot;/2020/12/15/become-a-data-engineer-on-a-shoestring.html&quot;&gt;here are some curated suggestions&lt;/a&gt;), bringing up the right amount of motivation to finish them can get really tough. Our bootcamp delivers &lt;a href=&quot;/2021/07/26/how-to-teach-data-engineering.html&quot;&gt;the structure, the methodology and the environment&lt;/a&gt; that makes you achieve your goals.&lt;/p&gt;
&lt;h4 &gt;2. Issues and frustration when building data infrastructures&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;So you are the data person who somehow ended up being the one building, improving and maintaining the data infrastructure at your company. Yet you lack the proper training to make informed decisions about tooling and optimisations? Or you are being held back from implementing your ideas, concepts and models due to the lack of data engineers or infra know-how? This is the thing that leads data analysts and data scientist leave their jobs and look for a more fulfilling path within the realm of data, which we know is more than achievable.&lt;/p&gt;
&lt;h4 &gt;3. Career improvement and salary increase&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Many data pros want to get more technical and earn the benefits of doing so (promotions, more job opportunities, higher salary etc.), and realise that &lt;a href=&quot;/curriculum/&quot;&gt;a compact yet intensive course would be the right path forward, covering the fundamentals&lt;/a&gt; that apply regardless of the industry they are in. When you are stuck at your job as BI analyst, data scientist or product person, improving your technical skillset delivers &lt;a href=&quot;/2020/09/22/data-engineer-salary-germany-2020.html&quot;&gt;a high return on investment&lt;/a&gt; even in the short term.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/learn_data_engineering_course_training_online.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;4. No fluff, no bs, just straight-talk&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Hard facts corroborated by benchmarks, hands-on experience with using tools and approaches at large organisations, and a contrarian yet truthful opinion are not easy assets to come by as a learner when players in the data ecosystem are hard-pressed to avoid them. Branded courses with marketing-as-a-curriculum (aka MaaC) will misguide learners and give them the warm and cozy feeling of certainty. &lt;em&gt;This is the best tool, stop asking questions.&lt;/em&gt; We just do the opposite — in all of our work.&lt;/p&gt;
&lt;h4 &gt;5. Data engineer job descriptions can be intimidating&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Have you ever read a job description for a data engineer role? Did it read like the line-up of an indie music festival (&lt;a href=&quot;https://greatexpectations.io/&quot;&gt;Great Expectations&lt;/a&gt;, &lt;a href=&quot;https://clickhouse.com/&quot;&gt;ClickHouse&lt;/a&gt; etc.), or the roster of an Italian football team (&lt;a href=&quot;https://luigi.readthedocs.io/en/stable/&quot;&gt;Luigi&lt;/a&gt;, &lt;a href=&quot;https://cassandra.apache.org/_/index.html&quot;&gt;Apache Cassandra&lt;/a&gt; etc.) or something Elon Musks PR team came up with as a grandiose joke (&lt;a href=&quot;https://www.terraform.io/&quot;&gt;Terraform&lt;/a&gt;, &lt;a href=&quot;https://aws.amazon.com/kinesis/data-firehose/&quot;&gt;Kinesis Firehose&lt;/a&gt; etc.)? If your answer is yes, I have news for you: this is the exact reason why even experienced data folks are often hesitant to apply for these positions. Part of the measurable transformation (before vs. after) that results from doing the course is that you&apos;ll feel empowered to apply, and you&apos;ll have the confidence to rock the interviews. &lt;a href=&quot;/2021/10/08/meet-the-graduates-michele-tassoni.html&quot;&gt;Right, Michele?&lt;/a&gt;&lt;/p&gt;
&lt;h4 &gt;6. Get a holistic &amp;amp; structured view on the data infrastructure ecosystem&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In order to get ready to start as data engineer after 12 weeks of training, you need to develop an understanding of the fundamental concepts, of the tooling landscape, of best practices, but also of the surrounding business context. Clouds and virtualisation? Real-life coding challenges from renowned organisations? Recommendations for diving deeper into ML, data modeling and dataops? It&apos;s all in our &lt;a href=&quot;/curriculum/&quot;&gt;curriculum&lt;/a&gt;.&lt;/p&gt;
&lt;h4 &gt;7. Small cohorts and personalised learning experience&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Some franchise bootcamps received significant backlash after filling their virtual classrooms with 30-40 students, which rendered organic teacher-learner and learner-learner interactions difficult. Active participation and asking questions should be encouraged all the time, this is how a coding bootcamp experience is supposed to be more than just looking at way too many talking heads in a Zoom window, where you are more or less just a number. Keeping the cohorts deliberately small allows us to focus on individual needs, customise career coaching and &lt;a href=&quot;https://www.coursereport.com/schools/pipeline-data-engineering-academy&quot;&gt;deliver an outstanding learning experience&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_online_course_bootcamp_best.jpg&quot; alt=&quot;Graffiti reading “believe in yourself” on a yellow wall&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;8. Expert teachers and renowned guest speakers&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Daniel and I have done our fair share of shenanigans in the realm of tech, data, education and building teams and digital products. We&apos;ve made loads of mistakes along the way, and we love to share our war stories so you don&apos;t end up repeating them. &lt;a href=&quot;/2021/02/15/learn-from-the-pros.html&quot;&gt;Our guest speakers&lt;/a&gt; share this attitude, and the opportunity to pick their brain is not something you&apos;ll get in any other school or conference.&lt;/p&gt;
&lt;h4 &gt;9. Future-proof knowledge&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Nobody really knows what the data infrastructure landscape is going to look like in 2030. Yet there are timeless best practices and tools attached to the smart data engineer&apos;s belt that allow for dealing with new and shiny trends of the day. There is a reason why we turn to SQL, Python and Makefiles in 2021, and there are reasons why we apply concepts like TCO before making decisions on tooling.&lt;/p&gt;
&lt;h4 &gt;10. Solving Real problems and building new Products is rewarding&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s not easy to stand out from the job seeking masses by doing the same pre-defined exercises as all the others. What would you say if I told you that the assignments at Pipeline Academy are designed in a technology-agnostic and solution-oriented fashion? We care about helping you figure things out rather than giving you &apos;&lt;em&gt;fill in the blanks&lt;/em&gt;&apos; type of assignments. If you like puzzles, you&apos;ll love the course.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - October 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/11/04/the-data-janitor-letters-october-2021.html"/>
   <updated>2021-11-04T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/11/04/the-data-janitor-letters-october-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.swyx.io/cloudflare-go/&quot; target=&quot;_blank&quot;&gt;Eating the Cloud from Outside In&lt;/a&gt;&lt;br&gt;&lt;em&gt;Shawn Wang, Developer Experience, &lt;/em&gt;&lt;a href=&quot;http://temporal.io/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Temporal.io&lt;/em&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;AWS is playing Chess. Cloudflare is playing Go.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/lightspeed-venture-partners/why-lightspeed-invested-in-clickhouse-a-database-built-for-speed-b67ec2d5f041&quot; target=&quot;_blank&quot;&gt;Why Lightspeed invested in ClickHouse: a database built for speed&lt;/a&gt;&lt;br&gt;&lt;em&gt;Gaurav Gupta, VC, Lightspeed Venture Partners&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;$250M Series B financing of ClickHouse.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://seattledataguy.substack.com/p/day-in-the-life-of-a-data-engineer&quot; target=&quot;_blank&quot;&gt;Day In The Life Of A Data Engineer — What Do Data Engineers Do?&lt;/a&gt;&lt;br&gt;&lt;em&gt;SeattleDataGuy&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;There’s no better time to jump into the world of data engineering.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.marksblogg.com/roapi-rust-data-api.html&quot; target=&quot;_blank&quot;&gt;ROAPI: An API Server for Static Datasets&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mark Litwintschik, #bigdata Consultant&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;ROAPI is an API Server that exposes CSV, JSON and Parquet files without the need to write any code.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.streamlit.io/announcing-streamlit-1-0/&quot; target=&quot;_blank&quot;&gt;Announcing Streamlit 1.0! 🎈&lt;/a&gt;&lt;br&gt;&lt;em&gt;Adrien Treuille, Co-Founder and CEO, Streamlit&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Streamlit used to be the simplest way to write data apps. Now it&apos;s the most powerful.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.timescale.com/blog/function-pipelines-building-functional-programming-into-postgresql-using-custom-operators/&quot; target=&quot;_blank&quot;&gt;Function pipelines&lt;/a&gt;&lt;br&gt;&lt;em&gt;David Kohn, Developer, Timescale&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Building functional programming into PostgreSQL using custom operators.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://aws.amazon.com/blogs/big-data/implement-a-slowly-changing-dimension-in-amazon-redshift/&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;I&lt;/strong&gt;mplement a slowly changing dimension in Amazon Redshift&lt;/a&gt;&lt;br&gt;&lt;em&gt;Milind Oke and Bhanu Pittampally, Amazon&lt;/em&gt;
&lt;/h4&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://engineering.hometogo.com/dbt-at-hometogo-ece067987267&quot; target=&quot;_blank&quot;&gt;DBT at HomeToGo. Creating a scalable framework&lt;/a&gt;&lt;br&gt;&lt;em&gt;Gijs de Kruif, HomeToGo&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Nothing new under the sun.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html&quot; target=&quot;_blank&quot;&gt;How to Execute Pandas Workloads in a Distributed Manner With Apache Spark&lt;/a&gt;&lt;br&gt;&lt;em&gt;Hyukjin Kwon and Xinrong Meng, Databricks&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If you have to do it you have to do it :facepalm:&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://cuddly-octo-palm-tree.com/posts/2021-10-31-better-bash-functions/&quot; target=&quot;_blank&quot;&gt;Bash functions are better than I thought&lt;/a&gt;&lt;br&gt;&lt;em&gt;Gary Verhaegen, Senior Software Engineer, Digital Asset&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Given all that, I simply do not understand why people keep recommending the {} syntax at all.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - September 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/10/15/the-data-janitor-letters-september-2021.html"/>
   <updated>2021-10-15T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/10/15/the-data-janitor-letters-september-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://stratechery.com/2021/cloudflares-disruption/&quot; target=&quot;_blank&quot;&gt;Cloudflare’s Disruption&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ben Thompson, Stratechery&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;S3’s margin is R2’s opportunity.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://matduggan.com/operations-is-not-developer-it/&quot; target=&quot;_blank&quot;&gt;Operations is not Developer IT&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mathew Duggan, DevOps Manager, GAN Integrity&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s not their fault, they were told this was easy.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://newsletter.pragmaticengineer.com/p/project-management-in-tech&quot; target=&quot;_blank&quot;&gt;How Big Tech Runs Tech Projects and the Curious Absence of Scrum&lt;/a&gt;&lt;br&gt;&lt;em&gt;Gergely Orosz&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A survey of how tech projects run across the industry highlights Scrum being absent from Big Tech. Why is this, and are there takeaways others should take note of?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://wraptext.equals.app/who-are-analysts-technically/&quot; target=&quot;_blank&quot;&gt;Who are analysts, technically?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Bobby Pinero, CEO and Co-Founder, Equals&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Rather than needing to be an impossible combination of statistician, developer, and business expert, analysts can simply be great critical thinkers.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://towardsdatascience.com/re-evaluating-kafka-issues-and-alternatives-for-real-time-395573418f27&quot; target=&quot;_blank&quot;&gt;Re-evaluating Kafka: issues and alternatives for real-time&lt;/a&gt;&lt;br&gt;&lt;em&gt;Olivia Iannone Technical Writer at Estuary&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Kafka’s challenges have exhausted many an engineer on the path to successful data streaming. What if there was an easier way?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://quix.ai/performance-limiations-python-client-libraries/&quot; target=&quot;_blank&quot;&gt;A very detailed comparison of Python stream processing libraries&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mike Rosam, Cofounder and CEO, Quix&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;“To successfully use Flink in production you must invest serious resources … estimate more than 18 months.”&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://eugeneyan.com/writing/first-rule-of-ml/&quot; target=&quot;_blank&quot;&gt;The First Rule of Machine Learning: Start without Machine Learning&lt;/a&gt;&lt;br&gt;&lt;em&gt;Eugene Yan, Applied Scientist, Amazon&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Having robust data pipelines and high-quality data labels also suggests you’re ready for machine learning.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://veekaybee.github.io/2021/09/23/enlightenment/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Reaching MLE (machine learning enlightenment)&lt;/em&gt;&lt;/a&gt;&lt;em&gt;&lt;br&gt;Vicki Boykis, Machine Learning Engineer, Automattic&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Once, on a crisp cloudless morning in early fall, a machine learning engineer left her home to seek the answers that she could not find, even in the newly-optimized Google results.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Meet The Graduates: Michele Tassoni</title>
   <link href="https://dataengineering.academy/2021/10/08/meet-the-graduates-michele-tassoni.html"/>
   <updated>2021-10-08T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/10/08/meet-the-graduates-michele-tassoni.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we’ll share some of the stories that Daniel and I get to watch unfold at Pipeline Academy. Check out what our graduates have to say about the course, how they’ve tackled its challenges and what they are doing now with their new data engineering superpowers.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Peter: Michele, it&apos;s great to see you again. Thanks for taking the time to have a chat with me. Can I ask you to please give us a short introduction about yourself?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Michele: My name is Michele. I&apos;m Italian, but I live in Germany for nine years. I am a biologist by education and, after my master thesis,  following my passion for technology I decided to continue my studies in Germany. Here I could learn about drug discovery, data analysis and image processing.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/michele_tassoni_data_engineer.jpg&quot; alt=&quot;Michele Tassoni&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Michele Tassoni, Data Engineer and graduate of Pipeline Academy’s founding cohort&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;This is where my story with data started, and I could see that I could learn and develop even further. Afterwards, I worked for two years at Wayfair as a data scientist. During this time I learned that building a model and making predictions is not enough in the industry. You need to be able to collect all your data, arrange your dataset, build new features, and only after that you can develop your model properly.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Furthermore, models need to go into production and for this I had to learn many new things. Because I like  to focus more on the technical aspects of the data processing, I decided to join Pipeline Academy.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: What were your expectations before you joined Pipeline Academy?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: Before joining Pipeline, I was looking for a bootcamp where I could grow my skills, and I wasn&apos;t sure whether to try to improve my data science skills or to go towards a more technical direction. Having talked to a couple of schools and their alumni, I thought that a data science course wouldn&apos;t be the best choice for me, because I wouldn&apos;t get that much out of it as from growing my engineering skills. Once I found out about Pipeline Academy I immediately thought that it would be a better option, and after talking to you I was almost immediately convinced to join.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: What did working through the twelve weeks feel like for you?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: As soon as I understood how the 12 weeks were structured and by &lt;a href=&quot;/curriculum/&quot;&gt;looking at the curriculum&lt;/a&gt; I thought that it would be quite difficult to keep up with it. But once we had to choose our portfolio project my goal was to improve it week by week, trying to implement the most important and interesting learnings that were covered during that particular week. That required quite a bit of effort from my side, also during my free time or even weekends sometimes. However, I enjoyed working on it.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: ...and finally you ended up with a very impressive capstone project. What was it about?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: I decided to create a product (&lt;a href=&quot;https://github.com/mrvaita/footballers_value&quot; target=&quot;&quot;&gt;link to GitHub&lt;/a&gt;) that was collecting data from &lt;a href=&quot;http://transfermarkt.com/&quot;&gt;transfermarkt.com&lt;/a&gt;, so about football players, mostly. I collected data about individual players (e.g. weight, height, market value, etc.) for the past 50 years for the top five European championships. The data was stored in a database, restructured and modelled. From this little data warehouse, I was extracting the data back and showing it on a dashboard via a web application. But there was a tiny trick about where this data warehouse lived.  I wanted to integrate &lt;a href=&quot;https://greatexpectations.io&quot; target=&quot;&quot;&gt;Great Expectations&lt;/a&gt; into my project, which was a bottleneck because my development environment became too large to be deployed on PythonAnywhere for example. So I ended up deploying it on my Raspberry Pi!&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Tell me about the most memorable moments of your time in the course. The good, the bad, the surprising moments, whatever comes top of your mind.&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: Oh, that&apos;s hard. I really enjoyed the guest speakers: for example &lt;a href=&quot;/2021/02/15/learn-from-the-pros.html&quot;&gt;Martin (Dr. Martin Loetzsch) for me was the best, followed by the &quot;Facebook guy&quot;, Bence and Daniel from Shopify&lt;/a&gt;. They were inspiring because listening to them I understood that to be good at something, a lot of effort is required.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Also, in my opinion if you join this bootcamp without a good base knowledge and the right level of motivation, it&apos;s easy to fall behind. That&apos;s why I would say miracles can happen in this school, yet a lot comes from you as a person and as a student, and your motivation and your willingness to push for those three months. It&apos;s really not that difficult to do it, and at the end of the course you will have learned a lot.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Did your perception of the data engineering role change during or after the cohort?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: So I have to say that probably without the school I wouldn&apos;t have applied for the data engineering position that I&apos;m currently having. I would have flagged this particular job as something that I wouldn&apos;t have been a good candidate for. But due to the course I was not scared to apply for it.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: If you would have to summarise, how did the bootcamp contribute to your success?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: When starting a new job, you have to learn the whole data architecture of your new company as quickly as possible and as well as you can.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;During the bootcamp, we covered many times concepts like how a data warehouse architecture looks like, how data is transferred from a backend database into the data warehouse and how the data moves from a warehouse to the business intelligence tool. This has helped me a lot during the first two months at my new job.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: This brings me to the question: what are you doing now?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: I am a Data Engineer at &lt;a href=&quot;https://www.contorion.de/&quot;&gt;Contorion&lt;/a&gt;. The team that I joined, maintains and improves the ETL and the data warehouse. This data is meant for our business intelligence and analytics teams who will then work with it developing new business insights. We also develop our own business intelligence tool. I&apos;m looking forward to the time when my team will start to develop new data products, and I can contribute to them. That will be amazing.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: If you could give advice to any prospective data engineer, what would it be?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;M: I always think that I never had a real good mentor, I always had to do everything by myself. But there were some moments during my studies where I could learn from more experienced people, and that always gave me such a boost. My advice is to try to find someone as your mentor who can lead you to become a better professional.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Thank you very much for the interview, Michele.&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - August 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/09/20/the-data-janitor-letters-august-2021.html"/>
   <updated>2021-09-20T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/09/20/the-data-janitor-letters-august-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@mrtrustworthy/from-data-driven-to-driving-data-the-dysfunctions-of-data-engineering-34c34496ed8e&quot; target=&quot;_blank&quot;&gt;From Data Driven to Driving Data — The dysfunctions of Data Engineering&lt;/a&gt;&lt;br&gt;&lt;em&gt;MrTrustworthy&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Many “data driven” initiatives are failing even though they had the best engineers on the task and picked the “best” stack of technologies.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://analyticsengineers.club/whats-an-olap-cube/&quot; target=&quot;_blank&quot;&gt;What&apos;s an OLAP cube? 🎲&lt;/a&gt;&lt;br&gt;&lt;em&gt;Claire Carroll, Analytics Engineer, analyticsengineers.club&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;OLAP cubes were this intimidating concept, and the more they read, the less they understood, but it turns out that they aren’t that confusing in practice. This is a trend that I’ve seen a lot in data engineering / data modeling where jargon is used as a gatekeeper.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://openlineage.io/blog/dataquality_expectations_facet/&quot; target=&quot;_blank&quot;&gt;Expecting Great Quality with OpenLineage Facets&lt;/a&gt;&lt;br&gt;&lt;em&gt;Michael Collado, Staff Software Engineer, Datakin&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Good data is paramount to making good decisions - but how can you trust the quality of your data and its dependencies?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://jlongster.com/future-sql-web&quot; target=&quot;_blank&quot;&gt;A future for SQL on the web&lt;/a&gt;&lt;br&gt;&lt;em&gt;James Long, Software Engineer, Stripe&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Absurd-sql is a persistent backend for SQLite on the web. That means it doesn’t have to load the whole db into memory, and writes persist.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.marksblogg.com/minio-aws-s3-hdfs.html&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;MinIO: A Bare Metal Drop-In for AWS S3&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Mark Litwintschik, Big Data Consultant&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;MinIO offers an S3 gateway service that can allow you to expose Hadoop&apos;s distributed file system (HDFS) with an AWS S3-compatible interface.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://ably.com/blog/no-we-dont-use-kubernetes&quot; target=&quot;_blank&quot;&gt;No, we don’t use Kubernetes&lt;/a&gt;&lt;br&gt;&lt;em&gt;Maik Zumstrull, Site Reliability Engineer, Ably&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It also doesn’t make sense for a lot of companies that are currently going all-in on Kubernetes.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://highscalability.com/blog/2021/8/2/evolution-of-search-engines-architecture-algolia-new-search.html&quot; target=&quot;_blank&quot;&gt;Evolution of search engines architecture&lt;/a&gt;&lt;br&gt;&lt;em&gt;Julien Lemoine, Co-founder &amp;amp; CTO, Algolia&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We look at some key milestones in the evolution of search engine architecture. We also describe the challenges those architectures face today.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - July 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/08/28/the-data-janitor-letters-july-2021.html"/>
   <updated>2021-08-28T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/08/28/the-data-janitor-letters-july-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://erikbern.com/2021/07/07/the-data-team-a-short-story.html&quot; target=&quot;_blank&quot;&gt;Building a data team at a mid-stage startup: a short story&lt;/a&gt;&lt;br&gt;&lt;em&gt;Erik Bernhardsson, Working on something, &quot;Bernco&quot;&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The data culture is driven both from above (the CEO pushing for it) as well as from below (people in the trenches). It&apos;s OK to fail if at least you learned something from it.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://benn.substack.com/p/analytics-is-at-a-crossroads&quot; target=&quot;_blank&quot;&gt;Analytics is at a crossroads&lt;/a&gt;&lt;br&gt;&lt;em&gt;Benn Stancil, Chief Analytics Officer + Founder, Mode&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The world is full of great analysts. Will we have the courage to go looking for them?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://scattered-thoughts.net/writing/against-sql/&quot; target=&quot;_blank&quot;&gt;Against SQL&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jamie Brandon, independent researcher&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;SQL is the only widely-used implementation of the relational model, and it is: inexpressive, incompressible, non-porous.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://pedram.substack.com/p/for-sql&quot; target=&quot;_blank&quot;&gt;For SQL&lt;/a&gt;&lt;br&gt;&lt;em&gt;Pedram Navid, Head of Data, Hightouch&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Why I&apos;m so over-protective of my data people.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://corecursive.com/066-sqlite-with-richard-hipp/&quot; target=&quot;_blank&quot;&gt;The Untold Story of SQLite With Richard Hipp&lt;/a&gt;&lt;br&gt;&lt;em&gt;Adam Gordon Bell, CoRecursive Podcast&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;On today’s show, I’m talking to Richard Hipp about surviving becoming core infrastructure for the world.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - June 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/08/02/the-data-janitor-letters-may-2021-june.html"/>
   <updated>2021-08-02T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/08/02/the-data-janitor-letters-may-2021-june.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.getdbt.com/analytics-engineering/&quot; target=&quot;_blank&quot;&gt;The Analytics Engineering Guide&lt;/a&gt;&lt;br&gt;&lt;em&gt;dbt Labs&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Collaborating as a data team to produce excellent datasets -- some parts are bullshit, but it&apos;s an interesting read.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.snowflake.com/blog/welcome-to-snowpark-new-data-programmability-for-the-data-cloud/&quot; target=&quot;_blank&quot;&gt;Welcome to Snowpark: New Data Programmability for the Data Cloud&lt;/a&gt;&lt;br&gt;&lt;em&gt;Isaac Kunen, Senior Product Manager, Snowflake&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Two words: Java functions.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://news.ycombinator.com/item?id=27582645&quot; target=&quot;_blank&quot;&gt;Accidentally exponential behavior in Spark&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ivan Vergiliev, Tech Lead, Leanplum&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Don&apos;t use Spark for tasks that require complex logic.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://veekaybee.github.io/2021/06/20/the-ritual-of-the-deploy/&quot; target=&quot;_blank&quot;&gt;The ritual of the deploy&lt;/a&gt;&lt;br&gt;&lt;em&gt;Vicki Boykis, Machine Learning Engineer, Tumblr&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Deploying is a ritual. It’s a sacred place, a quiet place, and a dangerous place, where anything can happen. In deployment, the system is in a fragile state, and you are in a fragile state.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.gentlydownthe.stream/&quot; target=&quot;_blank&quot;&gt;Gently Down the Stream&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mitch Seymour, Illustrator, Author, Founder, Round Robin Publishing&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A gentle introduction to Apache Kafka.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210621073654/https://technically.substack.com/p/whats-kafka-and-what-does-confluent&quot; target=&quot;_blank&quot;&gt;What&apos;s Kafka and what does Confluent do?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Justin Gage, Technically&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Help with solving Kafka-esque data problems&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://techcrunch.com/2021/06/01/cloudera-to-go-private-as-kkr-cdr-grab-it-for-5-3b/&quot; target=&quot;_blank&quot;&gt;Cloudera to go private as KKR &amp;amp; CD&amp;amp;R grab it for $5.3B&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ron Miller, TechCrunch&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Cloudera was once one of the hottest Hadoop startups, but over time the shine has come off that market, and today it went private.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;&lt;a href=&quot;https://about.gitlab.com/press/releases/2021-06-30-meltano-spins-out-of-gitlab-raises-seed-funding-led-by-gv.html&quot; target=&quot;_blank&quot;&gt;Meltano Spins Out of GitLab, Raises $4.2M in Seed Funding Led by GV to Enhance Open Source Data Integration&lt;/a&gt;&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&quot;Meltano aims to bring the entire data lifecycle into the DataOps Era.&quot; Wut?&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>How to Teach Data Engineering</title>
   <link href="https://dataengineering.academy/2021/07/26/how-to-teach-data-engineering.html"/>
   <updated>2021-07-26T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/07/26/how-to-teach-data-engineering.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Frequently we receive positive comments about &lt;a href=&quot;/curriculum/&quot;&gt;our curriculum&lt;/a&gt; and our &lt;a href=&quot;/overview/&quot;&gt;general approach&lt;/a&gt; to data engineering, especially in the light of the hype vs. the realness that surrounds it. However, the selected content would be ineffective and empty for our students without the right delivery. Selecting the fitting pedagogical approach and combining it with the most effective didactical methods are key to running a successful course, so here are the whys and hows behind what a Pipeline Academy student is dealing with at the bootcamp.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;As pointed out when talking &lt;a href=&quot;/2021/07/02/the-data-engineering-portfolio-project.html&quot;&gt;about data engineering portfolio projects&lt;/a&gt;, not having best practices in place forces you start from scratch and figure out yourself what would make sense in terms of teaching this constantly evolving subject matter. But what can go wrong — you ask?&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Well, it might happen that the learning process is simply ineffective, and you just don&apos;t get the gist.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You might leave a school and realize that whatever you&apos;ve learned is totally detached from real life, and you as a result fail to connect the dots and the &quot;career kickstart&quot; does not happen.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Potentially you could lose interest during the course and drop out.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Not being upfront and honest about the learning process and what kind of engagement and mental effort studying takes from students can be extremely misleading, and is something that we see a lot of coding bootcamps struggle with.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Making sure people can access the course fairly easily independently from their socioeconomic background is also something we talk about a lot internally.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;On top of this, data engineering is kind of a special case: you have to have strong principles and fundamentals to be able to tackle random incoming unstructured tasks and issues when working as a data engineer, and they do help when you&apos;re trying to maintain your sanity in general. This becomes very clear when you realise that a controlled environment is rarely something that these professionals deal with, and accepting this is an essential step in the right direction.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;So you need some experience and the attitude to get it right, and design a meaningful and engaging learning experience - &lt;a href=&quot;https://medium.com/edtechx360/the-need-for-pedagogy-market-fit-in-edtech-25d050418885&quot;&gt;a pedagogy market fit&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/teachin_data_engineering_course.jpg&quot; alt=&quot;A video call participant wearing a Chewbacca mask&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Daniel + Zoom + Chewbacca mask&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Not our first rodeo&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I have spent over seven years in the e-learning industry. I have  worked in higher education, in vocational education and have contributed to the digitalisation of the traditional school system&apos;s teaching methodologies. I&apos;ve presented conference papers on learning theory and practice at conferences like &lt;a href=&quot;https://oeb.global&quot;&gt;Online Educa Berlin&lt;/a&gt; (2004, 2005) and &lt;a href=&quot;https://microlearning.org&quot;&gt;Microlearning Innsbruck&lt;/a&gt; (2006) on connectivism, on edutainment and on social software like wikis.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The other co-founder and career coach at Pipeline Academy, Peter holds a double-degree in business administration and teaching, and he has taught marketing research at his alma mater while studying there. He has worked in various environments mentoring and supporting junior talent and helped navigate them through the realm of emerging technologies and the related competencies. He has been in charge of one of the world&apos;s top rated data science bootcamps and has coached various students transitioning to the field of data and supported them in the process of finding jobs and leveling up their careers.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I guess it&apos;s safe to say that we&apos;re professionally well-equipped to make the curriculum work for our students, but it&apos;s crucial to follow the latest research and digital tools in order to succeed on the long run.&lt;/p&gt;
&lt;h4 &gt;Pedagogical concept&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Contextual_learning&quot;&gt;Contextual learning&lt;/a&gt; is based on the &lt;a href=&quot;https://en.wikipedia.org/wiki/Constructivism_(philosophy_of_education)&quot;&gt;constructivist theory of teaching and learning&lt;/a&gt;. Learning takes place when teachers are able to present information in such a way that students are able to construct meaning based on their own experiences.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Contextual learning has the following characteristics:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;emphasizing problem solving,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;recognizing that teaching and learning need to occur in multiple contexts,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;assisting students in learning how to monitor their learning and thereby become self-regulated learners,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;anchoring teaching in the diverse life context of students,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;encouraging students to learn from each other,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;employing authentic assessment.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;As opposed to the more traditional &lt;em&gt;&quot;I&apos;ll show you how it&apos;s done, repeat after me, now you should be able to do it too&quot;&lt;/em&gt;-kind of teaching and learning environment — not that there is anything wrong with that when used in the appropriate context.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Why constructivism? The answer is easy: engineering is about solving problems. Data engineering is about solving (mostly) unstructured problems while filtering the massive hype around tooling and the ecosystem in general. There will be no handbooks, only a compass and a map. You have to build yourself an internal tool for navigating, while we help you sketch your interpretation of the world; that&apos;s the best you can do.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Every interaction (or the lack of it) permeates the curriculum with these paradigms. Assignments, teamwork and learning to fail and try again. That will be it. We&apos;ve tried and tested this approach: it&apos;s tough and it works.&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Trust the process.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Based on the experience of ourselves and our graduates we think these are the major considerations you should take when thinking about whether we are the right choice for you:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The bootcamp is hands-on, effort and outcome oriented.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You have to write code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;We will emulate real-life work scenarios.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You will solve problems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You have to take the initiative.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;We&apos;ll engage you in frontal instruction, solo assignments, research, working in pairs and teams, and you&apos;ll be presenting and explaining concepts and tools to others.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Standups, retros, open tickets, open Pull Requests are bread and butter for us.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You&apos;ll prepare for job interviews by answering practical questions and solving very real coding challenges.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Sync and async communication will happen all the time on Slack and Zoom.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You will report bugs with context, describe your attempts to resolve the situation, you will practice asking for help efficiently.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;This is not a simulation bubble, this is entropy itself.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Things will break as its their nature.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You might have to reinstall your computer once or twice.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;From time to time, you&apos;ll curse — and that&apos;s totally ok: &quot;This is not a pop album&quot; as Ice-T pointed out oh-so wisely in Warning.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/learn_data_engineering-kp6yqb.jpg&quot; alt=&quot;Yoda meme: much to learn you still have, my young padawan&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;If you are up for the fun, this is what you&apos;ll end up with when graduating from Pipeline Academy:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Confidence&lt;/strong&gt;: you&apos;ll realize that the invested amount of curious exploration, the effort put into trying again over and over, your continuous attempt at connecting the dots between what makes sense vs. what you are told, and the work spent collaborating with others correlates directly with your self-confidence for your job interviews and your future career. We provide the stage, but the show is about you.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Context&lt;/strong&gt;: are you intimidated when checking out data engineer job descriptions, because they are full of seemingly random tools you&apos;ve never heard of? Assessing and benchmarking them, using them and learning their quirks will enable you to deal with most of the data stacks out there, even the ones that don&apos;t exist yet.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Code you own&lt;/strong&gt;: as a data engineer, your code will most likely be looked at when applying for a job. Proving the right competencies and attitude with the work you&apos;ve delivered will feel rewarding and will make you stand out.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Engineering Portfolio Project</title>
   <link href="https://dataengineering.academy/2021/07/02/the-data-engineering-portfolio-project.html"/>
   <updated>2021-07-02T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/07/02/the-data-engineering-portfolio-project.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Being the co-founder of the pioneering data engineering bootcamp means having no real blueprint on how to do things. No franchise books, tutorials or blogposts to lean on. Having a decade of experience in data, and half of that on top in the e-learning industry I had to sit down and think again with the head of the hiring manager who had been interviewing data engineers in Berlin for 7 years now, condensing what I have seen in about a thousand interview sessions.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;If you join Pipeline Academy, it&apos;s very likely that you come to the bootcamp with the goal of leveling up in your career, meaning that you are going to have go through applications and job interviews after the course. In this process, you&apos;ll have to prove that you can execute the tasks listed in the job description, and that you are the right &lt;a href=&quot;/2021/05/05/5-specialisations-for-data-engineers.html&quot;&gt;Data Engineer/ML engineer/Data Product Owner/...&lt;/a&gt; for supporting the goals of your future employer.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The products and projects you&apos;ve built constitute a massive part of making a positive impression as an engineer, so let&apos;s take a look at what we consider fundamental expectations.&lt;/p&gt;
&lt;h4 &gt;How should a professional data engineering portfolio project (aka capstone project) look like?&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Peter and I start every month with &lt;a href=&quot;https://www.meetup.com/pipeline-data-engineering-academy-berlin/&quot;&gt;a virtual open house session&lt;/a&gt; where we introduce Pipeline Academy to all the interested folks who join in. We keep repeating like a mantra that our aim for our graduates is that they should leave the course with three things in their pockets:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Context&lt;/strong&gt;: i.e. understanding the ecosystem and the driving forces of data engineering,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Confidence&lt;/strong&gt;: for having the right attitude for solving unstructured problems,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Code they own&lt;/strong&gt;: everything they&apos;ve produced during the course.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Don&apos;t forget - and I am stating the obvious here - our expectations are broadly speaking not different from what startups and tech companies are looking for.&lt;/p&gt;
&lt;h4 &gt;Functionality&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Your portfolio project should have the functionality of a typical real world data stack. It does not have to be complex or complicated, straightforward is better. Matching your tooling decisions with the business circumstances shows you listen, think holistically and care about the &lt;em&gt;why&lt;/em&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Functionality that you should cover is:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;acquisition of data from a database, an API, a queue/tracker, a scraper,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;an automated ETL process, using a Prefect Cloud or a dbt cloud setup is a good choice - a cronjob and a Makefile is always a big plus for me,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;loading data to a datawarehouse - hints: SQLite, DuckDB can do wonders, and having data quality measures in place would not hurt either,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;if you do plain vanilla machine learning, you&apos;re already considered for a machine learning engineer role nowadays,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;it is deployed easily (Makefile, Dockerfile, something executable), maybe in the cloud even,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;your codebase has some CD/CI on it -- Github Actions will do, so you are dataops now,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;it serves data to humans (web interface) or machines (API).&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Subtle hint: you can see that &lt;a href=&quot;/curriculum/&quot;&gt;our curriculum&lt;/a&gt; is pretty much aimed at making sure you can do all of this by the end of the bootcamp course. If you have a data engineering portfolio project that fits this description I would happy to review it and have a chat about it with you.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_portfolio_capstone_project_machine_learning.png&quot; alt=&quot;Tweet by @svpino: pick a problem, build a model, wrap it in a simple API, host it, build a small app that uses it, set up a retraining pipeline, set up monitoring, write about all of this&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Source: &lt;/em&gt;&lt;a href=&quot;https://twitter.com/svpino/status/1392416514079350789&quot;&gt;&lt;em&gt;@svpino’s twitter feed&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Skills&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;These are the skills I&apos;d like to see demonstrated in a portfolio project:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you can explain your/the project&apos;s goal in simple terms (you have a one-pager README.md or something similar — &lt;a href=&quot;https://tom.preston-werner.com/2010/08/23/readme-driven-development.html&quot;&gt;Readme Driven Development&lt;/a&gt; can help here: describe in plain English what you&apos;re trying to achieve, describe what you don&apos;t know yet and still have to research),&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you have an architectural doodle/blueprint so one can grasp the components and their relationships,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;your code is okay enough to be actually read,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you demonstrate structured thinking, communication skills, understanding of the fundamental concepts, and collaboration capabilities (e.g. &lt;em&gt;&quot;I&apos;ve learned this from that blogpost and I opened an issue on the repo of this tool because I hit a wall&quot;&lt;/em&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;It does not have to be complicated -- it&apos;s better if it&apos;s not --, it does not have to solve all the problems of mankind. It just should deliver what you say it should.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s about showing &lt;em&gt;how&lt;/em&gt; you were thinking about a problem and &lt;em&gt;how&lt;/em&gt; you&apos;ve delivered an imperfect solution that is suitable for your constraints. It will have tradeoffs and that&apos;s totally fine, and I&apos;m happy to read about why and how you&apos;ve ended up making certain decisions.&lt;/p&gt;
&lt;h4 &gt;Interest and Attitude&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If you can deliver on the above, that&apos;s already a decent start in my book. However, you should not forget that the closer you get to the final stages of a job interview process, the more your future coworkers are going to scrutinise your soft skills and the so-called &quot;fit&quot;. In general, this is what a lot of people are looking for in a future coworker in engineering when it comes to general attitude:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you are confident in approaching and dealing with problems,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you understand K.I.S.S. as a principle,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you are determined to build a resilient system so you are not on call on the weekends for fixing bugs,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.clairecodes.com/blog/2019-05-15-cv-driven-development/&quot;&gt;you don&apos;t build things just because you can on the job&lt;/a&gt; — that is called a hobby,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you solve a problem without introducing a new one,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;tools are tools and not means,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you frequently ask questions, you listen to the answer and you think about them,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;you don&apos;t start a Spark cluster to load a CSV.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Do not underestimate the power of showing that you are somebody who others would like to sit next to. It makes a world of difference.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_portfolio_project_interview.jpg&quot; alt=&quot;The Office meme: Keep It Simple Stupid — great advice, hurts my feelings every time&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Examples&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Here are a couple of examples delivered by the students of Pipeline Academy. It should be noted that these products and capstone projects are being created under special circumstances (think time pressure, certain individual goals ingrained, focusing on solving one specific problem etc.).&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Tomek Florek: &lt;a href=&quot;https://github.com/xtomflo/zeit-ort&quot;&gt;Feature enrichment API for data science and analytics&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Amy Raygada: &lt;a href=&quot;https://github.com/amy6586/menu_generator&quot;&gt;Menu generator based on nutritional values&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Michele Tassoni: &lt;a href=&quot;https://github.com/mrvaita/footballers_value&quot;&gt;Data aggregation and data analysis app to determine the value of football players&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Michail Koskinas: &lt;a href=&quot;https://github.com/mkoskinas/class_to_hiphop&quot;&gt;A hip-hop playlist recommender based on the trackID of a classical music piece from Spotify&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Sujit Badle: &lt;a href=&quot;https://github.com/sb-olr/Crypto-data-ETL-backtest-Pipeline-project&quot;&gt;Crypto data ETL backtest&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;EDIT:  &lt;a href=&quot;/2022/02/18/data-engineering-capstone-project-data-product.html&quot;&gt;click here&lt;/a&gt; to see even more portfolio projects that our graduates have built during 2021.&lt;/p&gt;
&lt;h4 &gt;Don&apos;t forget&lt;/h4&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Tailor your capstone project according to your career goals: you can put an emphasis on code, architecture, communication, processes, quality etc. depending on what kind of role you are going for.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Make sure that your project is consistently presented and explained the right way, even for people who are not familiar with its context.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Practice explaining the why and the how: expect questions that uncover your train of thought.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Ask for feedback: wherever you have a chance, ask the hiring manager or whoever is reviewing your portfolio for feedback.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Pipeline Academy on the Data Engineering Podcast Vol.2</title>
   <link href="https://dataengineering.academy/2021/06/29/pipeline-academy-on-the-data-engineering-podcast-vol2.html"/>
   <updated>2021-06-29T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/06/29/pipeline-academy-on-the-data-engineering-podcast-vol2.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Daniel and I had the opportunity to sit down and have a chat again with Tobias Macey from the Data Engineering Podcast. His show is a mainstay in the data engineering community, and it was a pleasure to revisit the events of the last 8 months together with him. &lt;a href=&quot;https://www.dataengineeringpodcast.com/pipeline-data-engineering-academy-retrospective-episode-198/&quot;&gt;Listen to the episode right here!&lt;/a&gt;&lt;br&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;While &lt;a href=&quot;/2020/09/23/data-engineering-podcast-pipeline-academy.html&quot;&gt;Daniel’s first appearance on the show from September 2020&lt;/a&gt; was more about our very specific approach to solving engineering problems and turning this into a curriculum and a proper learning concept, and other future plans, this time we mostly speak in past tense.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We’ve talked about what launching Pipeline Academy felt like and we’ve shared some insights about the experiences of our founding cohort. &lt;a href=&quot;https://www.dataengineeringpodcast.com/pipeline-data-engineering-academy-retrospective-episode-198/&quot;&gt;Highlights, proud moments, challenges, lessons learned, takeaways for the future — you name it, it’s all in there&lt;/a&gt;. This episode is especially recommended for people who are considering joining a bootcamp and/or learning data engineering, and for the ones amongst you who are asking themselves why context, code and confidence are so important in order to get started in the universe of data and software engineering.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data-engineering-podcast-logo.jpg&quot; alt=&quot;Data Engineering Podcast logo&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;About the host and the show (taken from &lt;a href=&quot;https://web.archive.org/web/20200930155323/https://www.dataengineeringpodcast.com/about/&quot;&gt;the website of the podcast&lt;/a&gt;):&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;“Tobias Macey is a dedicated engineer with experience spanning many years and even more domains. He currently manages and leads the Technical Operations team at MIT Open Learning where he designs and builds cloud infrastructure to power online access to education for the global MIT community. He also owns and operates &lt;/em&gt;&lt;a href=&quot;https://www.boundlessnotions.com/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Boundless Notions, LLC&lt;/em&gt;&lt;/a&gt;&lt;em&gt; where he offers design, review, and implementation advice on data infrastructure and cloud automation.”&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;“The Data Engineering Podcast tackles a new approach to data management every week. Each new episode provides useful and informative insights into the projects, platforms, and practices that data engineers, team leaders, and data scientists need to know about to learn and grow in their career.”&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Like and subscribe on &lt;a href=&quot;https://podcasts.apple.com/us/podcast/data-engineering-podcast/id1193040557?mt=2&quot;&gt;Apple Podcasts&lt;/a&gt; and give the show a follow on &lt;a href=&quot;https://twitter.com/DataEngPodcast&quot;&gt;twitter&lt;/a&gt;!&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - May 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/06/24/the-data-janitor-letters-may-2021.html"/>
   <updated>2021-06-24T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/06/24/the-data-janitor-letters-may-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://jasnonaz.medium.com/analytics-engineering-everywhere-d56f363da625&quot; target=&quot;_blank&quot;&gt;Analytics Engineering Everywhere&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jason Ganz, Cofounder, ChattyKathi&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Or why in five years every organization will have an Analytics Engineering team.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.metabase.com/learn/data-diet/analytics/data-model-mistakes&quot; target=&quot;_blank&quot;&gt;Common data model mistakes made by startups&lt;/a&gt;&lt;br&gt;&lt;em&gt;Metabase&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;It’s important to note that the anti-patterns we’ll discuss below are specific to startups.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://miles2code.com/data/datawarehouse/2021/05/11/data-modeling-principles.html&quot; target=&quot;_blank&quot;&gt;16 fundamental principles for transforming data in a warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Rahul Jain, Head of BI and Data Engineering, Beat&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The current discourse on data can get a little tiring because of its over focus on tooling.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210625002703/https://www.narrator.ai/blog/using-postgresql-as-a-data-warehouse/&quot; target=&quot;_blank&quot;&gt;Using PostgreSQL as a Data Warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Cedric Dussud, Cofounder, &lt;/em&gt;&lt;a href=&quot;http://narrator.ai/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Narrator.ai&lt;/em&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;With some tweaking Postgres can be a great data warehouse. Here&apos;s how to configure it.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.linkedin.com/pulse/why-spark-right-tool-etl-work-kevin-bair/&quot; target=&quot;_blank&quot;&gt;Why Spark is NOT the right tool for ETL work&lt;/a&gt;&lt;br&gt;&lt;em&gt;Kevin Bair, Director Sales Engineering, Snowflake&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;My larger point here is when you have a hammer (spark) everything looks like nail.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20230117064812/https://towardsdatascience.com/using-apache-airflow-dockeroperator-with-docker-compose-57d0217c8219&quot; target=&quot;_blank&quot;&gt;Using Apache Airflow DockerOperator with Docker Compose&lt;/a&gt;&lt;br&gt;&lt;em&gt;Flávio Clésio, Staff Engineer Data/Machine Learning, Artsy&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I personally believe that Airflow + Docker it’s a good combination for flexible, scalable, and hassle-free environments for ELT/ETL tasks.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://eugeneyan.com/writing/machine-learning-metagame/&quot; target=&quot;_blank&quot;&gt;The Metagame of Applying Machine Learning&lt;/a&gt;&lt;br&gt;&lt;em&gt;Eugene Yan, Applied Scientist, Amazon&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;When designing systems, less is more.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://blog.cleverelephant.ca/2021/05/postgis-20-years.html&quot; target=&quot;_blank&quot;&gt;PostGIS at 20, The Beginning&lt;/a&gt;&lt;br&gt;&lt;em&gt;Paul Ramsey, Executive Geospatial Engineer, Crunchy Data&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;All the development was done on the trusty Sun Ultra 10 I had taken out a $10,000 loan to purchase when starting up the company.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;&lt;a href=&quot;https://www.reddit.com/r/ExperiencedDevs/comments/nmodyl/drunk_post_things_ive_learned_as_a_sr_engineer/&quot; target=&quot;_blank&quot;&gt;Drunk Post: Things I&apos;ve learned as a Sr Engineer&lt;/a&gt;&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;SQL is king. Airflow is shit, yes.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data and sustainability: Interview with Zsofia Bognar From Ecosia</title>
   <link href="https://dataengineering.academy/2021/06/08/data-sustainability-zsofia-bognar-from-ecosia.html"/>
   <updated>2021-06-08T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/06/08/data-sustainability-zsofia-bognar-from-ecosia.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Pipeline Academy works with a number of selected organisations and supports them with in-house data engineering trainings and with guidance related to their data infrastructure. One of these companies is &lt;a href=&quot;https://www.ecosia.org/&quot;&gt;Ecosia&lt;/a&gt;, the search engine that plants trees. The (growing) data team of the Berlin-based startup is led by Zsofia Bognar, and we’ve had the chance to interview her about the challenges at the intersection of data and sustainability.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/large_zsofia_web_81162f0272.jpg&quot; alt=&quot;Zsófia Bognár&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Zsofia Bognar, BI &amp;amp; Data Team Lead at Ecosia (&lt;a href=&quot;https://www.linkedin.com/in/zsofiabognar/&quot;&gt;LinkedIn profile&lt;/a&gt;)&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;&lt;strong&gt;Peter: What is your professional background and what drove you towards a career in data?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zsofia: I’ve been working in the data field for the last 12 years so it’s hard for me to imagine life WITHOUT data :) . But jokes aside, I think what drew me to work with data most is how by analysing data we can get closer to understanding the world around us. I studied Economic Sciences, a part of which was statistics. For a lot of people statistics is a really dry discipline, but I found it fascinating how with the help of statistical methods we can gain fresh insights in a concise way for just about anything. &lt;br&gt;&lt;br&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: You are currently in charge of the data team at Ecosia. What kind of value is generated from user data for the organisation and what kind of team is dedicated to data initiatives?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: At Ecosia we are on a mission to build a digital green companion that plants billions of trees. We want to help our users  make sustainable choices, and be a role model in the transition to a sustainable society. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;As a data team we contribute to this mission by delivering data products in three main areas: &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;1) We focus on helping product teams make data-informed choices in line with their KPIs and their own team mission. We help them monitor and interpret their results. Advising and educating teams plays a big role in ensuring team success. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;2) We create reports to evaluate different marketing efforts to understand what the most informative way to reach our users is. We not only want to attract new users but make actionable information available about the climate emergency in an approachable way and we as data team help the marketing team to identify these areas. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;3) We build capabilities for the tree team to evaluate  their tree planting projects, which in turn also delivers value to our tree planting partners since our tree planting officers are able to communicate transparently and quickly the impact and possible areas of improvement for each project. Our goal is to know about the state of every single tree we help to plant.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Currently we are a four person team: Arnaud, our tech lead, is a true data pipeline plumber who constantly improves our setup and finds better solutions. Nikki, our Product/Marketing Analyst, is the jack-of-all-trades in our team in its best meaning; knowing the company and business model inside out, identifying problems and working towards solving them.  And we have our latest joiner, Elise, who is a really great data storyteller. She understands the needs of the stakeholders and how information can be meaningfully visualised for them&lt;br&gt;&lt;br&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: What are your teams biggest technical and organisational challenges?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: We are a small team with a lot of stakeholders: currently we have 8 stakeholder teams. This means in the past year we had to completely change our mindset of what a data team is. Previously we operated as a service team where we had ad-hoc requests from the teams. With the resource constraints our backlog grew, the team became frustrated and our stakeholders became under-served. Balancing our legacy infrastructure and soaring stakeholder needs was a real struggle. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We understood this was not sustainable so we changed to a data product team mindset. Now we identify the product needs with user stories, and build our products around this. We introduced agile methods and work in sprints. We also educate our users on how best to use these products and integrate their feedback on the data UX design. We have seen that this has increased the velocity of the team but also data has integrated much better into the day-to-day of Ecosia. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Our main technical challenge is handling truly big data sets. There are almost no BI or ETL products or solutions that we can use out of the box. This means we are regularly dealing with challenges to optimise solutions for our needs along the entire data pipeline. For this reason, it’s hard for us to find partners who can usefully consult us on these challenges, which is why it was great for us to work with Pipeline Academy.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/ecosia_pipeline_academy.png&quot; alt=&quot;Ecosia search page: every search removes 1kg of CO2, 127,231,714 trees planted&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Ecosia.org&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: Ecosia acts very transparently when it comes to financial information and respects the privacy of users. How is this translated into the day-to-day operations of a data team?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: We are being transparent both inside the team and towards our stakeholders as well. We have a data-as-code infrastructure so everything we produce is transparent and documented. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We are also showing things honestly as they are. We don’t have user groups with different levels of data access, everyone in Ecosia can see everyone else’s dashboards. We find this very effective as well as there is always some cross-over of which data is needed for the different teams and sharing them openly increases trust in the data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: If you would start out in data today, what kind of skills would you focus on in the first place?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: I think understanding where one’s own strengths lie and how they fit into the data products that teams deliver. There are so many aspects of data products and all of them are crucial in producing great output. From understanding stakeholder needs to planning, execution, building the pipeline, designing the UX for users and actually understanding the trends. To master them all is almost impossible, so understanding your own growth path and fitting it together with the team needs is essential. Having good skills in SQL and python helps a lot though, and understanding tools that became industry standards like dbt or Airflow gives one a good advantage.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: You&apos;ve hired Pipeline Academy to support your team&apos;s move towards analytics engineering, a hot topic in the world of BI. What kind of outcomes do you expect from the upskilling training?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: For us as a small and high velocity team it has been essential that we make the right tooling decisions and avoid some of the pitfalls those tools provide. So for us, trial and error is rarely a path we can take. We started to work together with Pipeline Academy because we knew Daniel and Peter have extremely in-depth knowledge of the tooling landscape and are also really embedded in the data community, so they can really well advise us on datatools across the entire pipeline. I also like how they didn’t offer us a solution as the Holy Grail but gave us an evaluation framework by which we can make these decisions ourselves.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;P: I understand your team is growing, and you are looking for data engineering talent. What kind of hard and soft skills do you think makes a newcomer in data stand out?&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Zs: We need &lt;a href=&quot;https://explore.ecosia.org/careers-data-engineer&quot;&gt;someone who understands the principles of data infrastructure as code and can help us to be even more reliable towards our users&lt;/a&gt;. We’d like to have someone with the mindset who sees data products in the holistic way we do in the team and doesn’t just work on engineering tickets. We would like to have someone who is excited about learning but also can mentor others in the team. It would  be great to have someone who has experience with Airflow, dbt and with  different cloud providers. &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - April 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/05/31/the-data-janitor-letters-april-2021.html"/>
   <updated>2021-05-31T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/05/31/the-data-janitor-letters-april-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.allthingsdistributed.com/2021/04/s3-strong-consistency.html&quot; target=&quot;_blank&quot;&gt;Diving Deep on S3 Consistency&lt;/a&gt;&lt;br&gt;&lt;em&gt;Dr. Werner Vogels, CTO, &lt;/em&gt;&lt;a href=&quot;http://amazon.com/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Amazon.com&lt;/em&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We built S3 on the design principals that we called out when we launched the service in 2006, and every time we review a design for a new feature or microservice in S3, we go back to these same principles.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://calpaterson.com/metadata.html&quot; target=&quot;_blank&quot;&gt;&lt;strong&gt;We were promised Strong AI, but instead we got metadata analysis&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Cal Paterson, contract software engineer&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;How simple structured data trumps clever machine learning.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210409152141/https://towardsdatascience.com/a-comprehensive-framework-for-data-quality-management-b110a0465e83&quot; target=&quot;_blank&quot;&gt;A Comprehensive Framework for Data Quality Management&lt;/a&gt;&lt;br&gt;&lt;em&gt;Chau Vinh Loi, Data Scientist, ANZ Australia&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;How to monitor and maintain Data Quality to make sure the data meets certain standards for specific business use-cases&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://mitchellsilv-79772.medium.com/layering-your-data-warehouse-f3da41a337e5&quot; target=&quot;_blank&quot;&gt;Layering Your Data Warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mitchell Silverman, Analytics Engineer, Spotify&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I never thought I would be comparing my work in data engineering to the great Mike Myers but Ogres and Data Warehouses have a lot in common. Both are misunderstood by most and both can save the day when called upon.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://benn.substack.com/p/metrics-layer&quot; target=&quot;_blank&quot;&gt;The missing piece of the modern data stack&lt;/a&gt;&lt;br&gt;&lt;em&gt;Benn Stancil, Chief Analytics Officer + Founder, Mode&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The core problem is that there’s no central repository for defining a metric.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/explorium-ai/benchmarking-sql-engines-for-data-serving-prestodb-trino-and-redshift-1c5f16d6e5da&quot; target=&quot;_blank&quot;&gt;Benchmarking SQL engines for Data Serving: PrestoDb, Trino, and Redshift&lt;/a&gt;&lt;br&gt;&lt;em&gt;Anton Peniaziev, data and machine learning engineer, Explorium&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data serving is a special business case, which demands real-time low-latency on small queries, the ability to scale and withstand abrupt peak loads, and high concurrency.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://towardsdatascience.com/introducing-a-dataflow-management-system-backed-up-by-prefect-aws-and-github-actions-3f8c0eef2eb2&quot; target=&quot;_blank&quot;&gt;Introducing a Dataflow Management System Backed Up by Prefect, AWS, and Github Actions&lt;/a&gt;&lt;br&gt;&lt;em&gt;Maikel Penz, Senior Data Engineer, Spidertracks&lt;/em&gt;
&lt;/h4&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://erikbern.com/2021/04/19/software-infrastructure-2.0-a-wishlist.html&quot; target=&quot;_blank&quot;&gt;Software infrastructure 2.0: a wishlist&lt;/a&gt;&lt;br&gt;&lt;em&gt;Erik Bernhardsson, Ex-CTO, Better&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I mean, as a user, I can set up a static website in AWS, but it takes 45 steps in the console and 12 of them are highly confusing if you never did it before.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://mtlynch.io/litestream/&quot; target=&quot;_blank&quot;&gt;How Litestream Eliminated My Database Server for $0.03/month&lt;/a&gt;&lt;br&gt;&lt;em&gt;Michael Lynch, Builder of @TinyPilotKVM&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data persistence for people who hate database servers.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://speakerdeck.com/charity/cd&quot; target=&quot;_blank&quot;&gt;It is time to fulfill the promise of CI/CD&lt;/a&gt;&lt;br&gt;&lt;em&gt;Charity Majors, Cofounder/CTO, @honeycombio&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Why your software should be auto-deployed within 15 minutes after you merge it, with no manual gates. This is the key to high performing teams and high-quality software.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Top 5 Specialisations for Data Engineers</title>
   <link href="https://dataengineering.academy/2021/05/05/5-specialisations-for-data-engineers.html"/>
   <updated>2021-05-05T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/05/05/5-specialisations-for-data-engineers.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;When mankind embarks on a journey of exploring a new subject matter, the first part of the ride is usually bumpy and full of ideas and initiatives that turn out to be erroneous in hindsight. It is very true that failing is essential for learning - just ask the people who call themselves &quot;&lt;em&gt;serial entrepreneur&lt;/em&gt;&quot; on LinkedIn. When working on figuring out how to approach a problem we tend to systematically reinvent and restructure hypotheses, tools, methodologies and processes over and over again in order to maximise the likelihood of a desired outcome, and later on we try our best to optimise all this for efficiency. This is an established form of discovery that aims at generating value, all in the name of &lt;span style=&quot;text-decoration:line-through&quot;&gt;capitali...&lt;/span&gt; prosperity.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Just take the evolution of the automobile or the continuous development of modern medicine as an example. I would argue, that we as a society are just getting started with finding out how software development and maintenance should be sustainably institutionalised, and the growing pains of today can be vividly felt when talking to people who have spent some time working in tech. Setting up organisational structures, distributing tasks, managing workloads, streamlining communication, exploring new methods of collaboration are all popular subjects for bestsellers for a reason. As a consequence of the quickly evolving consumer expectation towards the products and services built upon software, we keep iterating and we remain on the lookout for somebody with answers - or at least with a totemic acronym we can put on pink post-it notes.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But within this circus, there are smaller stages with plays of comparable conflicts: the stage of data being one of them. The quickly changing nature of the tools and processes applied to generate value out of the growing amount of available data requires constant reiteration of the skillset of professionals working with it, and a continuous realignment of organisational roles and responsibilities. You can multiply this uncertainty-soaked complexity with demographic, cultural, geographic and organisational factors that all have an undeniable influence on the daily business of a team working with technology.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This environment and the constant change is what makes it so difficult to create blueprints for long-term career pathways within data, therefore everyone who is planning on staying en vogue should care to remain ahead of the curve... somehow.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_specialisations.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;how to Level up&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The established, basic trifecta of data roles in 2021 consists of the data analyst, the data scientist and the data engineer - yet a lot is happening when you look at further career development. The most straightforward way of increasing the value one generates is to take on more decision making responsibility (financial responsibility, personnel responsibility) - think Senior Data Engineer, VP Data Engineering, Head of Data etc. But there are more and more specialised roles popping up that mix the fundamental competencies of a data engineer with different group of expertise to create a powerful combination. These are some of the roles that I see are highly sought after on the job market, these skill-combos are carrying loads of value to businesses who are ultimately looking for casting the right actors to star the show.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Interestingly, we can observe the participants of the first cohort of Pipeline Academy already getting a feel for their own personal paths and how they are starting working towards that goal. We are putting a lot of emphasis on &lt;a href=&quot;/curriculum/&quot;&gt;the fundamentals of data engineering&lt;/a&gt; at the bootcamp, yet ultimately it&apos;s the combination with the unique professional backgrounds and personal ambitions of our students that make them really stand out on the job market.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Let&apos;s take a look at five selected roles that are worth a deeper dive.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/analytics_engineer_bi_engineer.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;1. Analytics Engineer or BI Engineer&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The role of the Analytics Engineer or Business Intelligence Engineer started popping up in 2018 in more and more blogs and articles mainly thanks to this thing called &lt;a href=&quot;https://web.archive.org/web/20210419095229/https://blog.getdbt.com/what-is-an-analytics-engineer/&quot;&gt;dbt&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Today, if you’re a “modern data team” your first data hire will be someone who ends up owning the entire data stack.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;This title describes basically &lt;a href=&quot;https://web.archive.org/web/20210121111745/https://www.hashpath.com/2020/12/an-analytics-engineer-is-really-just-a-pissed-off-data-analyst/&quot;&gt;a pissed off data analyst&lt;/a&gt; who is by default able to generate insights out of data, and with some added engineering skills they become an end-to-end powerhouse for setting up solid BI pipelines without the help of an additional engineer or software developer. This idea is part of a broader trend of enabling data consumers to access what they want without intermediaries. Analytics engineers are highly valued especially in the days where hiring data engineers is fairly difficult due to a lack of available supply. Some are even going as far as saying that &quot;&lt;a href=&quot;https://quantumblack.medium.com/data-engineerings-role-is-scaling-beyond-scope-and-that-should-be-celebrated-ca9fa1cb8cbb&quot;&gt;80% of analytics is effectively data engineering&lt;/a&gt;&quot;... Definitely one of the most promising roles you can enter in 2021. There are a handful of forward-thingking scale-ups who are already hiring for this role in Berlin.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Competencies: analytical domain knowledge, SQL, Python and ETL/ELT knowledge, dbt, data visualisation and storytelling skills.&lt;/p&gt;
&lt;h4 &gt;2. Machine Learning Engineer or AI Engineer&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;ML engineers are data scientists who can productionize their models in order to solve business challenges - this might sound like an oversimplified definition, yet &lt;a href=&quot;https://cloud.google.com/certification/machine-learning-engineer&quot;&gt;that&apos;s pretty much it&lt;/a&gt;. Does this really get you into one of the best paid positions within data? Well, kind of. Currently I see two directions to arrive at this role: either you come from a data science angle and learn how to set up architectures and integrate sophisticated models into the equation in a robust manner (notice the DS —&amp;gt; DE pivot), or you come from the computer science/software engineer/data engineer realm and you level up your understanding of the mathematical/computational models you are supposed to build or paste into your pipelines. &lt;a href=&quot;https://www.reworked.co/information-management/what-do-machine-learning-engineers-do-anyway/&quot;&gt;The key expectations include&lt;/a&gt;:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&quot;&lt;em&gt;Translating the work of data scientists from environments such as Python/R notebooks analytics applications.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Creation of web services/APIs for serving ML/AI model results and enabling access to customers or internal teams.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Automating model training and evaluation processes.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Automating feature engineering, ensuring data for model training is cleaned and readily available and facilitation of the flow of data between ML/AI models and an organization’s data systems.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;And the list goes on. You see, putting something &apos;into production&apos; is what does the trick, this is where the business value is generated. &lt;a href=&quot;https://web.archive.org/web/20200719064232/https://towardsdatascience.com/ml-engineers-are-losing-their-jobs-learn-ml-anyway-87e19523cd9b&quot;&gt;Some critics even go as far as saying&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;Classic algorithms + domain knowledge + niche datasets are going to solve most real problems, not deep neural nets. Most of us aren’t working on self-driving cars.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;... and there is certainly a lot of truth to that.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Competencies: experience with the fundamentals of machine learning models, computer science or software engineering background combined with the latest data engineering expertise, SQL, Python and the command line, concepts like Continuous Delivery and Continuous Integration, understanding of Docker and maybe Kubernetes.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/dataops_machine_learning_ops.jpg&quot; alt=&quot;Workers joining a section of large pipeline in a trench&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;3. DataOps and Machine Learning Ops&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data architectures need maintenance, sometimes significantly more than traditional or legacy software stacks to perform and keep the output quality at a high level at all stages of the lifecycle. The reason for this are the constantly changing tools and business requirements that demand attention, and of course &lt;a href=&quot;https://www.ibm.com/blogs/journey-to-ai/2019/12/the-difference-between-dataops-and-devops-and-other-emerging-technology-practices/&quot;&gt;the pipelines that need fixing as a result&lt;/a&gt;. Shouldn&apos;t this be done by the data engineers, you ask? Well, not really, but you&apos;ll definitely hear about small-scale data teams that don&apos;t cultivate ops in general. Mature (or bold) data organisations who have identified proper machine learning usecases with a positive ROI are the ones who first have to start thinking about maintaining the systems their ML engineers have put in place as well. &lt;a href=&quot;https://ml-ops.org/&quot;&gt;Running the operations-side of data&lt;/a&gt; is often considered the least glorious part of the job, yet it&apos;s a very sought-after and rewarding role.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Competencies: SQL, Python, command line, understanding algorithmic complexity, operations experience, knowledge of concepts like CD/CI, logging, monitoring and profiling, understanding of Docker and  Kubernetes.&lt;/p&gt;
&lt;h4 &gt;4. Data Product Manager or Data Product Owner&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In order to successfully lead a cross-functional team one requires not only a feel for prioritising and managing multiple tasks in parallel, but also excellent communication and social skills. But as we dive deeper and deeper into the dark seas of technical complexity when designing and deploying software and data products, navigating becomes exponentially tougher. If companies would like to meet consumer needs and deliver the right features in an efficient and agile manner, the people in charge of the product roadmap need to be aware of the high-level technical requirements and consequences of their decisions (think data infrastructure costs, maintenance effort, data governance and quality across the org, stakeholder expectation management etc.). It is becoming clear that while the field of building websites or mobile apps is experiencing more and more commoditisation and is therefore turning simpler and more accessible, &lt;a href=&quot;https://medium.com/analytics-and-data/new-roles-of-analytics-the-data-product-owner-analytics-translator-c04bf0eacdad#:~:text=Data%2DScience%20Product%20Owner&amp;amp;text=Overall%20the%20company%20and%20team,of%20a%20data%20product%20owner.&amp;amp;text=Product%20Owners%20are%20meant%20to,prioritize%20the%20feature%20for%20development.&quot;&gt;dealing with the whole lifecycle of data products is much more complicated and requires and increased awareness of the latest tooling landscape, the methodologies and the lingo in data&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Data engineers who would like to be in charge of a product are well suited to fulfil this role, at the same time I am witnessing how experienced product people are leveling up their game to become data/AI POs by learning the essentials of data engineering at the bootcamp. Data leaders of the future, now is the time to do a deep dive into data infra; and in case ML comes into play in the consumer facing product, &lt;a href=&quot;https://medium.com/@adnanboz/product-managers-skills-for-survival-in-the-ai-era-3f59186a68d&quot;&gt;there are plenty of accessible online courses available for getting the knowledge you need&lt;/a&gt;. As Adnan Boz (founder of the AI Product Management Institute) &lt;a href=&quot;https://web.archive.org/web/20210420141818/http://thenextweb.com/news/ai-product-owners-needed-not-data-scientists-syndication&quot;&gt;puts it&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Since the development lifecycle of AI projects is based on “searching” rather than “planning”, companies need professionals who are trained to look at products as optimization problems rather than a programming problems.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Competencies: organisational skills, communication skills, understanding of the architecture and deployment process of data products and their consequences for the business and the user, stakeholder management.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_product_manager_owner_data_governance_czar.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;5. Data Governance Manager or Data Czar&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Talking to data professionals who are tired of hearing platitudes about a company being &quot;data-driven&quot;, a shared pattern seems to surface: data governance is not a nice-to-have anymore, but essential for utilising your data and your data staff. Now, what they are talking about is not the superficial &quot;oh, sure, we more or less understand GDPR and we do have a data warehouse&quot;-type of governance, but &lt;a href=&quot;https://tdan.com/complete-set-of-data-governance-roles-responsibilities/21589&quot;&gt;rules that are transparently applied across the whole organisation&lt;/a&gt;. The tendency of giving this role to one dedicated person or team who oversees and manages all potential conflicts and helps resolve them within the organisation is becoming more and more popular, especially as the legal responsibility of dealing with data is growing (and so are potential fines for companies). &lt;a href=&quot;https://web.archive.org/web/20210418002731/https://edx.readthedocs.io/projects/devdata/en/stable/internal_data_formats/data_czar.html&quot;&gt;Data czars manage data access for users&lt;/a&gt;, deal with data lineage, data quality, data dictionaries, support feature development with guidelines about collecting, storing, leveraging and sharing data in and outside of the company. However they are not to be confused with the data security manager (the Germans have a word for it: Datenschutzbeuftragte).&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This role will become a mainstay: take the increasing complexity of data management in general, mix it with the continuously evolving legal environment and add the growing demand of generating more and more value for shareholders through using data. It&apos;s a very fine line between becoming an overprotective naysayer versus being a thoughtful enabler of data teams.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Competencies: understanding data privacy and security best practices and legal requirements, stakeholder management skills, database management, encryption and decryption techniques, data quality measurement, documentation and advocacy.&lt;/p&gt;
&lt;h4 &gt;+1&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;...and something I did not mention above, although it is a real point of differentiation from the employer perspective: industry-specific expertise aka domain knowledge is a major asset. Understanding the mechanisms of a vertical or the language used in a certain sector will always bring you bonus points when applying to a job. If you leverage that the right way it can secure you a head start in the battle for a position.&lt;/p&gt;
&lt;h4 &gt;The only constant in life is change&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The lines between these roles are blurry at best, and they will undergo significant changes in the coming years. Engineering skills within data remain predominant however, they are kind of a superpower that can be combined with a surprising amount of unrelated competencies.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I guess the &lt;a href=&quot;https://web.archive.org/web/20200618170929/https://towardsdatascience.com/machine-learning-engineers-will-not-exist-in-10-years-c9cbbf4472f3&quot;&gt;main takeaway&lt;/a&gt; should be:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Most employers expect you to have overlapping skillsets. I feel like in the end it’s not about who gets wiped out, it’s about who is versatile enough to constantly adapt to the ever changing industry.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;If you are working on the frontlines of technology, you have to face the fact that this simply is the nature of this very quickly evolving environment, and you have to accept that being successful requires you to be curious and determined. Indeed adaptation is the magic word and since we’re far from being done exploring technology as a method, we can expect more change to come.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;At the same, this challenge is what makes data and tech such an exciting and rewarding place to be in.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;PS: In the next episode of Career Specialisations for Data Engineers: Sustainable Data Architect, Data Quality Engineer etc.&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;PPS: If you think I&apos;ve missed something, or you have suggestions for the above - feel free to shoot me an &lt;/em&gt;&lt;a href=&quot;mailto:info@dataengineering.academy&quot;&gt;&lt;em&gt;email&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - March 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/04/28/the-data-janitor-letters-march-2021.html"/>
   <updated>2021-04-28T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/04/28/the-data-janitor-letters-march-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://altinity.com/blog/clickhouse-nails-cost-efficiency-challenge-against-druid-rockset&quot; target=&quot;_blank&quot;&gt;Star Schema Benchmark: ClickHouse Nails Cost-Efficiency Challenge Against Druid &amp;amp; Rockset&lt;/a&gt;&lt;br&gt;&lt;em&gt;Alexander Zaitsev, Co-Founder &amp;amp; CTO, Altinity&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;When we use the same schema approaches as Druid and Rockset, ClickHouse significantly outperformed both whilst using the cheaper AWS setup.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.foxhound.systems/blog/sql-performance-with-union/&quot; target=&quot;_blank&quot;&gt;Speeding up SQL queries by orders of magnitude using UNION&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ben Levy and Christian Charukiewicz, Partners and Principal Software Engineers, Foxhound Systems&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;SQL’s UNION operation is not usually thought of as a means to boost performance.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://usefathom.com/blog/worlds-fastest-analytics&quot; target=&quot;_blank&quot;&gt;Building the world&apos;s fastest website analytics&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jack Ellis, Co-founder, Fathom&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Performing a migration is such a high adrenaline, stressful task.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://antonz.org/sqlite-is-not-a-toy-database/&quot; target=&quot;_blank&quot;&gt;SQLite is not a toy database&lt;/a&gt;&lt;br&gt;&lt;em&gt;Anton Zhiyanov&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Whether you are a developer, data analyst, QA engineer, DevOps person, or product manager - SQLite is a perfect tool for you.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/tiqets-tech/taming-the-dependency-hell-with-dbt-2491771a11be&quot; target=&quot;_blank&quot;&gt;Taming the Dependency Hell with dbt&lt;/a&gt;&lt;br&gt;&lt;em&gt;Rafael Barbosa, Data Engineering Team Lead, WeTransfer&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We’re using dbt to simplify the management of the views and build more trust in the data we store in our data warehouse.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/adaltas/storage-size-and-generation-time-in-popular-file-formats-48a23190c1da&quot; target=&quot;_blank&quot;&gt;Storage size and generation time in popular file formats&lt;/a&gt;&lt;br&gt;&lt;em&gt;Barthelemy Ngom, Solution Architect &amp;amp; Data Engineer, Adaltas&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;For archiving it is preferable to choose column based format and ORC.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210425180821/https://launchyourapp.meezeeworkouts.com/2021/03/why-we-dont-use-docker-we-dont-need-it.html&quot; target=&quot;_blank&quot;&gt;Why We Don’t Use Docker (We Don’t Need It)&lt;/a&gt;&lt;em&gt;&lt;br&gt;Nicky Rees, MeeZeeCo&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We get a single binary.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #15: Adam Vincze</title>
   <link href="https://dataengineering.academy/idataengineer/2021/04/14/idataengineer-confessions-interview-015.html"/>
   <updated>2021-04-14T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/04/14/idataengineer-confessions-interview-015.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;  &amp;lt;br/&amp;gt;  &amp;lt;iframe src=&quot;//www.youtube.com/embed/hQmj51LQlYo?wmode=opaque&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen&amp;gt;&amp;lt;/iframe&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/15_adam_wide.jpg&quot; alt=&quot;Adam, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #14: Alex Angelini</title>
   <link href="https://dataengineering.academy/idataengineer/2021/03/31/idataengineer-confessions-interview-014.html"/>
   <updated>2021-03-31T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/03/31/idataengineer-confessions-interview-014.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;  &amp;lt;br/&amp;gt;  &amp;lt;iframe src=&quot;//www.youtube.com/embed/FfT-aJ3CLsc?wmode=opaque&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen&amp;gt;&amp;lt;/iframe&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/14_alex_wide.jpg&quot; alt=&quot;Alex, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - February 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/03/30/the-data-janitor-letters-february-2021.html"/>
   <updated>2021-03-30T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/03/30/the-data-janitor-letters-february-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/h4&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://litestream.io/blog/why-i-built-litestream/&quot; target=&quot;_blank&quot;&gt;Why I Built Litestream&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ben Johnson, Builder, Litestream&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Despite an exponential increase in computing power, our applications require more machines than ever because of architectural decisions made 25 years ago. You can eliminate much of your complexity and cost by using SQLite &amp;amp; Litestream for your production applications.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.theregister.com/2021/02/25/google_kubernetes_autopilot/&quot; target=&quot;_blank&quot;&gt;Google admits Kubernetes container tech is so complex, it&apos;s had to roll out an Autopilot feature to do it all for you&lt;/a&gt;&lt;br&gt;&lt;em&gt;Tim Anderson, Senior Reporter, The Register&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;More expensive, less flexible, but easier and safer to use.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
Serverless Data Pipelines Made Easy with Prefect and AWS ECS Fargate&lt;br&gt;&lt;em&gt;Anna Anisienia, Python Engineer, TrailStone Group&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The easiest way to orchestrate your Python data pipelines.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.startdataengineering.com/post/apache-superset-tutorial/&quot; target=&quot;_blank&quot;&gt;Apache Superset Tutorial&lt;/a&gt;&lt;br&gt;&lt;em&gt;Joseph Machado, Senior Data Engineer, Narrativ&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In this post we go over the architecture of Apache Superset, connect to a warehouse and learn how to build charts and dashboards.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://skamille.medium.com/make-boring-plans-9438ce5cb053&quot; target=&quot;_blank&quot;&gt;Make Boring Plans&lt;/a&gt;&lt;br&gt;&lt;em&gt;Camille Fournier, Head of Platform Engineering, Two Sigma&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;You must do the work to go beyond vision, create concrete actions, and make boring plans.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210222111551/https://www.theolognion.com/hacker-hacks-into-startup-rewrites-complex-backend-with-postgresql-and-cron-forces-ceo-to-fire-40-of-devs-and-devops/&quot; target=&quot;_blank&quot;&gt;Hacker hacks into startup, rewrites complex backend with PostgreSQL and cron, forces CEO to fire 40% of devs and devops&lt;/a&gt;&lt;br&gt;&lt;em&gt;Matthew Solenya, Editor in Chief&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Alarmed, other startups in the industry are hiring more security specialists in order to reduce the risk of so-called &quot;green-hat hacker&quot; attacks.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://vole.wtf/kilobytes-gambit/&quot; target=&quot;_blank&quot;&gt;The Kilobyte’s Gambit ♟️💾 1k chess game&lt;/a&gt;&lt;br&gt;&lt;em&gt;Matt Round, Óscar Toledo G., Pinot W. Ichwandardi&lt;/em&gt;
&lt;/h4&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #13: András Dötsch</title>
   <link href="https://dataengineering.academy/idataengineer/2021/03/17/idataengineer-confessions-interview-013.html"/>
   <updated>2021-03-17T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/03/17/idataengineer-confessions-interview-013.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;  &amp;lt;br/&amp;gt;  &amp;lt;iframe src=&quot;//www.youtube.com/embed/duXWVSwqTuk?wmode=opaque&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen&amp;gt;&amp;lt;/iframe&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;br/&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/13_andras_wide.jpg&quot; alt=&quot;András, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Campus</title>
   <link href="https://dataengineering.academy/2021/03/10/the-campus.html"/>
   <updated>2021-03-10T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/03/10/the-campus.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;So many of you have been asking us about our new campus... nah, I&apos;m kidding, nobody is really thinking about classrooms or offices nowadays, because we tend to assume that every form of interpersonal interaction has been banned into the online space (and to our homes at the same time). Yet, as a firm believer in the superior value of in-person education I am very excited to walk you through our new location and tell you about what you should expect when joining Pipeline Academy.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Writing about the place we’ve picked for our bootcamp feels somehow like &lt;a href=&quot;https://youtu.be/zNtKT9_1KXQ&quot;&gt;a strange episode of MTV Cribs&lt;/a&gt;: the camera is too shaky, the tour feels like an exercise in humblebragging, and ultimately I&apos;m just showing you stuff I don&apos;t own. Nevertheless…&lt;/p&gt;
&lt;h4 &gt;Let me give you the tour: The Kiez&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Let&apos;s start with the area, which is favourite to a lot of people in Berlin. The campus is located on the western end of &lt;a href=&quot;https://de.wikipedia.org/wiki/Berlin-Kreuzberg&quot;&gt;Kreuzberg&lt;/a&gt; at Gneisenaustraße 27, close to Mehringdamm. This neighbourhood is defined by the variety of cozy parks, friendly streets and various international restaurants, which turns going out for lunch from an ordeal to a real highlight every single day. Historic spots like Gräfekiez, Viktoriapark, Tempelhofer Feld are just a few walking minutes away, and the ever-changing Bergmannstraße with the Marheineke Markthalle is even closer offering more options for grabbing groceries and food. (If you&apos;d like to get a vivid feel for how this Kiez has been evolving, go check out &lt;a href=&quot;https://www.tagesspiegel.de/berlin/bezirke/friedrichshain-kreuzberg/kreuzberg-ende-der-50er-jahre-bergmannkiez-wie-haste-dir-veraendert/9408936.html&quot;&gt;Gerd Nowakowski&apos;s lovely piece in the Tagesspiegel about it - in German&lt;/a&gt;.)&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Getting there should not be a problem for anyone, there are plenty of designated bike lanes in the surrounding area and public transport options are also plentiful. The closest stop right across the street is U7 Gneisenaustraße, but you can also take the U6 to Mehringdamm or the buses 140 and M19. This means that wherever you are commuting from within the city, it should be easy and hassle-free.&lt;/p&gt;
&lt;/div&gt;
 
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;come on in: The campus&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The area has a very rich history, and most of the surrounding buildings that survived the war are under heritage preservation. &lt;a href=&quot;https://de.wikipedia.org/wiki/Gneisenaustra%C3%9Fe#:~:text=Der%20Wohn%2D%20und%20Fabrikkomplex%20in,Lampenfabrik%20(alle%201943).):&quot;&gt;This is what went down behind these walls in the past 120 years&lt;/a&gt;: &lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;Der Wohn- und Fabrikkomplex in der Gneisenaustraße 27 wurde im Jahr 1900 errichtet. In früheren Jahren befanden sich hier unter anderem kleinere und größere Fabriken wie eine Messinglinienfabrik und Schriftgießerei, die Armaturen-Apparate-Fabrik Preschona von A. Meyer oder eine Lampenfabrik (alle 1943). Von 1949 bis 2013 diente das fünfgeschossige Gebäude als Produktionsstätte der Firma ROKA Robert Karst GmbH &amp;amp; Co. KG, die Steckverbindungen für die Autoindustrie fertigte, die danach einen neuen Standort bezog.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;The building with the old facade is &lt;a href=&quot;https://www.google.com/maps/@52.4911541,13.3957824,3a,75y,224.31h,110.46t/data=!3m6!1e1!3m4!1sbbGX-JhPqph6mNb1g3cU9w!2e0!7i13312!8i6656!5m1!1e2?hl=en&quot;&gt;still visible on Google Street View&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The campus has been established 2015 when the generous renovation of the magnificent old industrial building has been completed. Our partner and host, the &lt;a href=&quot;https://www.ciee.org/go-abroad/college-study-abroad/programs/germany/berlin/open-campus-block&quot;&gt;CIEE Global Institute&lt;/a&gt; has transformed the building into a &quot;&lt;a href=&quot;https://www.archdaily.com/772841/g27-ciee-global-institute-macro-sea#&quot;&gt;full-fledged “vertical campus” where living, academic classes, dining and socializing co-exist, creating a state-of-the-art interdisciplinary space. Functioning as a community incubator, the integration of the campus encourages learning both inside and outside of the classroom and facilitates connections that are vital to open exchange and debate.&lt;/a&gt;&quot; &lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;  sqs-gallery-container
  sqs-gallery-block-slideshow
  sqs-gallery-has-controls
  sqs-gallery-has-thumbnails
    sqs-gallery-block-show-meta sqs-gallery-block-meta-position-bottom
  sqs-gallery-block-show-meta
  block-animation-none
  clear&quot;&gt;
  &lt;div class=&quot;sqs-gallery&quot;&gt;            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_4.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_4.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_17.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_17.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_29.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_29.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_19.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_19.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_16.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_16.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_1.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_1.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
            &lt;div class=&quot;slide content-fill&quot; data-type=&quot;image&quot; data-click-through-url=&quot;&quot;&gt;                &lt;noscript&gt;&lt;img src=&quot;/images/posts/g27_macro_sea_3.jpg&quot; alt=&quot;&quot;&gt;&lt;/noscript&gt;
&lt;img src=&quot;/images/posts/g27_macro_sea_3.jpg&quot; alt=&quot;&quot;&gt;
              &lt;div class=&quot;color-overlay&quot;&gt;&lt;/div&gt;            &lt;/div&gt;
  &lt;/div&gt;
    &lt;div class=&quot;sqs-gallery-meta-container&quot;&gt;        &lt;div class=&quot;sqs-gallery-controls&quot;&gt;          &lt;a tabindex=&quot;0&quot; role=&quot;button&quot; class=&quot;previous&quot; aria-label=&quot;Previous Slide&quot;&gt;&lt;/a&gt;
          &lt;a tabindex=&quot;0&quot; role=&quot;button&quot; class=&quot;next&quot; aria-label=&quot;Next Slide&quot;&gt;&lt;/a&gt;
        &lt;/div&gt;
    &lt;/div&gt; &lt;!-- END .sqs-gallery-meta-container --&gt;
&lt;/div&gt;
    &lt;div class=&quot;sqs-gallery-thumbnails&quot;&gt;          &lt;img src=&quot;/images/posts/g27_macro_sea_4.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_17.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_29.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_19.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_16.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_1.jpg&quot; alt=&quot;&quot;&gt;
          &lt;img src=&quot;/images/posts/g27_macro_sea_3.jpg&quot; alt=&quot;&quot;&gt;
    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Here is the part where I gladly refer you to &lt;a href=&quot;https://www.archdaily.com/772841/g27-ciee-global-institute-macro-sea#&quot;&gt;Chris Mosier&apos;s photographs on ArchDaily&lt;/a&gt;, the photos above are their courtesy. I think they give you a better idea about how it feels to hang out on campus than I could ever do it justice.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The generous dorm on the site is available for our future students, so in case you don&apos;t have a place to stay in Berlin and would like to avoid the potential pitfalls of finding an apartment for the three months of the bootcamp, you can get an affordable all-inclusive student apartment at a really cool location. Just &lt;a href=&quot;/apply/&quot;&gt;hit us up&lt;/a&gt; for more info.&lt;/p&gt;
&lt;h4 &gt;Where the magic happens: The classroom&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Our classroom offers a versatile and flexible learning environment for our students, it allows us to adjust the setting according to the different requirements defined by the changing didactic methods we apply throughout the 12-weeks of &lt;a href=&quot;/curriculum/&quot;&gt;the data engineering bootcamp program&lt;/a&gt;. You can find us behind the bright yellow doors right next to the green Ampelmann.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/img_6900.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Although the classroom is being used as a streaming studio at the moment for delivering a professional online experience to our students (I am writing this March 10th, 2021), we have the option to  host in-person 1-on-1 sessions with our participants who ask for extra support. Both Daniel and I are eagerly waiting to have the whole class around and to finally working together in a non-virtual setting, but health and safety come first.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Exciting times are ahead of us, and we’re glad to have found a place that supports the launch of our small yet ambitious endeavour of building up Pipeline Academy. This is the part when the tour ends, but instead of saying goodbye, I’d like to take the opportunity to invite you to say hi if you are around!&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Don&apos;t forget, if you&apos;d like to get serious about a newly evolved profession that is &lt;a href=&quot;/2020/11/26/data-trend-report-prediction-for-2021.html&quot;&gt;likely to become an essential mainstay of the digital/data/innovation economy&lt;/a&gt;, an enabler of &lt;a href=&quot;/2021/01/10/sustainable-data-engineering.html&quot;&gt;sustainable digital product development&lt;/a&gt;, you can always join our data engineering training with a &lt;a href=&quot;/2021/01/10/bildungsgutschein-ready.html&quot;&gt;Bildungsgutschein from the AfA&lt;/a&gt; basically free of charge.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #12: Mikalai Syty</title>
   <link href="https://dataengineering.academy/idataengineer/2021/03/03/idataengineer-confessions-interview-012.html"/>
   <updated>2021-03-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/03/03/idataengineer-confessions-interview-012.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/KDThCw_nbVY?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/12_mikalai_wide_cover.jpg&quot; alt=&quot;Mikalai, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - January 2021</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/02/23/the-data-janitor-letters-january-2021.html"/>
   <updated>2021-02-23T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/02/23/the-data-janitor-letters-january-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.mihaileric.com/posts/we-need-data-engineers-not-data-scientists/&quot; target=&quot;_blank&quot;&gt;We Don&apos;t Need Data Scientists, We Need Data Engineers&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Mihail Eric, Machine Learning Scientist, Amazon Alexa AI&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;There are 70% more open roles at companies in data engineering as compared to data science.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210118001205/https://blog.arctype.com/sql-50-years/&quot; target=&quot;_blank&quot;&gt;What can we learn from SQL&apos;s 50 year reign? A story of 2 Turing Awards&lt;/a&gt;&lt;br&gt;&lt;em&gt;Felix Schildorfer, Chief Data Scientist, First Retail Inc.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The relational data model was introduced in 1970 and has dominated for 50 years. What led to its success? Building on first principles and Bushnell&apos;s law.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210224202829/https://blog.seekwell.io/gpt3&quot; target=&quot;_blank&quot;&gt;Automating my job by using GPT-3 to generate database-ready SQL to answer business questions&lt;/a&gt;&lt;br&gt;&lt;em&gt;Brian Kane, Data Engineer, SeekWell&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Now, I&apos;ve got a GPT-3 instance that takes a plain English question and translates it to SQL that really works on my database.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.dataengineeringpodcast.com/shopify-data-warehouse-with-dbt-episode-171/&quot;&gt;How Shopify Is Building Their Production Data Warehouse Using DBT - Episode 171&lt;/a&gt;&lt;br&gt;&lt;em&gt;Michelle Ark + Zeeshan Qureshi, Senior Data Engineer + Tech Lead/Engineering Manager, Shopify&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Structure the project to allow for multiple teams to collaborate in a scalable manner, have the additional tooling to address the edge cases, and the optimize the continuous integration process to provide fast feedback and reduce costs.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://github.com/oleg-agapov/data-engineering-book/blob/master/book/2-beginner-path/2-1-databases/databases.md&quot; target=&quot;_blank&quot;&gt;Introduction to Databases for Data Engineers&lt;/a&gt;&lt;br&gt;&lt;em&gt;Oleg Agapov, Data analyst/BI, &lt;/em&gt;&lt;a href=&quot;http://gog.com/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;GOG.com&lt;/em&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data engineers need to have a broad knowledge about ways of storing and processing data. Most of this knowledge will come with practice. But is still important to understand general ideas behind all concepts I&apos;ve explained here.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://davidxiang.com/2021/01/10/kafka-as-a-database/&quot; target=&quot;_blank&quot;&gt;Kafka As A Database? Yes Or No – A Summary Of Both Sides&lt;/a&gt;&lt;br&gt;&lt;em&gt;David Xiang, Engineering Team Manager, Squarespace&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I personally have never used a Kafka log as the source-of-truth for my data. Software development is hard enough as it is, even when trying to go “by the book.”&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tailscale.com/blog/an-unlikely-database-migration/&quot; target=&quot;_blank&quot;&gt;An unlikely database migration&lt;/a&gt;&lt;br&gt;&lt;em&gt;Brad Fitzpatrick + David Crawshaw, Late Stage Co-Founder + CTO, Tailscale&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The goal is to keep development speed as close to the early days of JSONMutexDB, when you could recompile and run locally in a fraction of a second and deploy ten times a day.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-threadripper-pro-workstation/&quot; target=&quot;_blank&quot;&gt;Achieving 11M IOPS &amp;amp; 66 GB/s IO on a Single ThreadRipper Workstation&lt;/a&gt;&lt;br&gt;&lt;em&gt;Tanel Põder, Co-founder, Gluent Inc.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Modern disks are so fast that system performance bottleneck shifts to RAM access and CPU.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #11: Micha Kunze</title>
   <link href="https://dataengineering.academy/idataengineer/2021/02/17/idataengineer-confessions-interview-011.html"/>
   <updated>2021-02-17T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/02/17/idataengineer-confessions-interview-011.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/l0YaBeAH2AA?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/11_micha_wide_cover.jpg&quot; alt=&quot;Micha, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Learn from the Pros: Guest speakers of our first cohort</title>
   <link href="https://dataengineering.academy/2021/02/15/learn-from-the-pros.html"/>
   <updated>2021-02-15T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/02/15/learn-from-the-pros.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;While following the core &lt;a href=&quot;/curriculum/&quot;&gt;12-week curriculum&lt;/a&gt; of the data engineering bootcamp, (almost) every Thursday a selected guest speaker will share their experience on the latest domain that’s being covered, and they will be available for AMA for our students: team structures, tech stacks, hiring processes and the future of the trade - no stone will remain unturned.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Data engineering feels like a black box for a lot of people even after spending some time with the subject matter, and I have to admit, it can be difficult to imagine the everyday life of a data professional. That is why we’ve invited some experienced practitioners to our (currently virtual) campus: our participants will have an opportunity to learn how they can put their knowledge to use, and see how others have built their careers on competences in the intersection of data and software engineering.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We’ll have freelancers who are enjoying working remotely, industry veterans who are on the lookout for new talent for their teams, subject matter experts who have designed our built significant parts of the digital products all of us use on a daily basis… and so on. It’s a very diverse list of perspectives.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Without any further ado, please check out the list of guest speakers for our very first cohort (in order of proposed appearance):&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/balazs-krich.jpg&quot; alt=&quot;Balazs Krich&quot;&gt;
&lt;h4 &gt;Balázs Krich&lt;br&gt;Lisbon&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Balázs is a data professional with exposure to product, data science and engineering. Having worked in media, banking, real estate, supply chain and e-commerce he also contributed to non-profit projects and data journalism as an Aaron Swartz Fellow at the Open Society Archives. A recent project of his covered collecting and standardizing all transactions in the European Strategic and Investment Funds (ESIF) between 2014 and 2020.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/balazs-krich/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://github.com/balkey/eu1420&quot;&gt;GitHub&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/bence-faludi.jpg&quot; alt=&quot;Bence Faludi&quot;&gt;
&lt;h4 &gt;Bence Faludi, Singapore&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Bence is a Data Engineer at Facebook. He made major contributions to several open source data munging tools like mETL, hamustro and night-shift. A recent conference talk of his is &lt;a href=&quot;https://www.youtube.com/watch?v=ArzohefZLE4&quot;&gt;Data Architecture 101 for Your Business&lt;/a&gt;. He was &lt;a href=&quot;https://www.youtube.com/watch?v=zkWuWWolvE4&quot;&gt;featured in the legendary #idataengineer podcast&lt;/a&gt;. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/bfaludi/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://github.com/bfaludi&quot;&gt;GitHub&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/dr.-martin-loetzsch.jpg&quot; alt=&quot;Dr Martin Loetzsch&quot;&gt;
&lt;h4 &gt;Dr. Martin Loetzsch,&lt;br&gt;Berlin&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Martin is the Chief Data Officer at &lt;a href=&quot;https://www.project-a.com/&quot;&gt;Project A Ventures&lt;/a&gt;, the operational VC. He built several data teams, and made major contributions to the open source ETL tool, mara. Two highlights of his recent conference talks are &lt;a href=&quot;https://www.youtube.com/watch?v=GdtFuOah-5c&quot;&gt;Data Warehousing with Python&lt;/a&gt; and &lt;a href=&quot;https://www.youtube.com/watch?v=whwNi21jAm4&quot;&gt;ETL Patterns with Postgres&lt;/a&gt;. He was also &lt;a href=&quot;https://www.youtube.com/watch?v=x_v0yYZQL48&amp;t=46s&quot;&gt;featured in our #idataengineer podcast&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/martin-loetzsch/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;http://www.martin-loetzsch.de&quot;&gt;Website&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/stefan-urbanek.jpg&quot; alt=&quot;Stefan Urbanek&quot;&gt;
&lt;h4 &gt;Stefan Urbanek, Taipei&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Stefan is a Data Infrastructure Architect who contributed to the infra of Facebook, Squarespace and Orange. His open source projects include the OLAP tool Cubes and ETL tool Bubbles. He talks about Cubes at &lt;a href=&quot;https://www.youtube.com/watch?v=-FDTK80zsXc&quot;&gt;Data Warehouse and Conceptual Modeling with Cubes 1.0&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/stefanurbanek/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://www.stiivi.com&quot;&gt;Website&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/miklos-koren.jpg&quot; alt=&quot;Miklós Koren&quot;&gt;
&lt;h4 &gt;Miklós Koren, Budapest&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Miklós teaches reproducible coding practices to economists to help them maximize their scientific impact. He believes in the command line, plain text, and that every problem can be solved with the right combination of Stata, Python, Julia, git, and make. He is Professor of Economics at &lt;a href=&quot;https://www.ceu.edu/&quot;&gt;Central European University&lt;/a&gt; and the Data Editor of the &lt;a href=&quot;https://www.restud.com/&quot;&gt;Review of Economic Studies&lt;/a&gt;, a leading scientific journal. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/miklos-koren-b535a42a/%20&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://koren.mk&quot;&gt;Research&lt;/a&gt;, &lt;a href=&quot;https://web.archive.org/web/20220419193422/https://learn.codedthinking.com/&quot;&gt;Training&lt;/a&gt;,  &lt;a href=&quot;https://twitter.com/korenmiklos&quot;&gt;Twitter&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/andreas-dewes.jpg&quot; alt=&quot;Andreas Dewes&quot;&gt;
&lt;h4 &gt;Andreas Dewes, Berlin&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Andreas is a serial entrepreneur whose current company, &lt;a href=&quot;https://kiprotect.com/&quot;&gt;KIProtect&lt;/a&gt; makes data security and data privacy easy for companies and organizations.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/andreas-dewes-39a53b13/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://github.com/KIProtect&quot;&gt;GitHub&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/rahul-jain.jpg&quot; alt=&quot;Rahul Jain&quot;&gt;
&lt;h4 &gt;Rahul Jain,&lt;br&gt;Berlin&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Rahul is a Principal Engineering Manager of Business Intelligence at &lt;a href=&quot;https://www.omio.com/&quot;&gt;Omio&lt;/a&gt; with decades of experience both as a manager and engineer.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/rahulj51/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://miles2code.com/&quot;&gt;Blog&lt;/a&gt;, &lt;a href=&quot;https://web.archive.org/web/20210227042823/https://www.mentoring-club.com/the-mentors/rahul-jain&quot;&gt;Mentoring Club Profile&lt;/a&gt;, &lt;a href=&quot;https://twitter.com/rahulj51&quot;&gt;Twitter&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/daniel-oliveira-filho.jpg&quot; alt=&quot;Daniel Oliveira Filho&quot;&gt;
&lt;h4 &gt;Daniel Oliveira Filho,&lt;br&gt;Berlin&lt;/h4&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Daniel is a Lead Production Engineer at &lt;a href=&quot;https://www.shopify.com/&quot;&gt;Shopify&lt;/a&gt;. He is a DevOps expert and practitioner who is also interested in the scaling questions of the online gaming industry. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/daniel-oliveira-filho-51045a21/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://github.com/Octops&quot;&gt;GitHub&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;img src=&quot;/images/posts/fl-vio-cl-sio.jpg&quot; alt=&quot;Flávio Clésio&quot;&gt;
&lt;h3 &gt;Flávio Clésio, Berlin&lt;/h3&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Flávio is a Staff Data Engineer at &lt;a href=&quot;https://www.artsy.net/&quot;&gt;Artsy&lt;/a&gt;. A machine learning and data engineer with teaching experience in data warehousing. A recent conference talk of his is &lt;a href=&quot;https://www.youtube.com/watch?v=8-eYUBgqOfM&quot;&gt;Preventing Revenue Leakage and Monitoring Distributed Systems&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/in/flavioclesio/&quot;&gt;LinkedIn&lt;/a&gt;, &lt;a href=&quot;https://flavioclesio.com&quot;&gt;Website&lt;/a&gt;﻿&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #10: Dr. Martin Loetzsch</title>
   <link href="https://dataengineering.academy/idataengineer/2021/02/03/idataengineer-confessions-interview-0010.html"/>
   <updated>2021-02-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/02/03/idataengineer-confessions-interview-0010.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/x_v0yYZQL48?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/10_martin_wide_cover.jpg&quot; alt=&quot;Martin, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>2020 - A Year To Remember</title>
   <link href="https://dataengineering.academy/2021/01/28/2020-a-year-to-remember.html"/>
   <updated>2021-01-28T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/01/28/2020-a-year-to-remember.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Some of you might think we are writing this blog just for you, our dearest readers, but the truth is that it mostly serves documentation purposes so in 10 or 20 years we can look back, reminisce and wonder how young and naive we were about starting a coding school in 2020. This post is basically a collection of events of the first 8 months of Pipeline Academy that Daniel and I consider significant in some way, so if you are new to our journey this might be a good starting point.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We started working on Pipeline Academy in March 2020. One could say our timing was not optimal, and I would wholeheartedly agree. As for many, the last year was about figuring out ways to make our business work &lt;em&gt;despite&lt;/em&gt; of all the mess that was happening in the realm of the global economy, and our private lives. But it&apos;s all about attitude... they say.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_the_office_2020.gif&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Right at the very start  we&apos;ve organised a mini online-bootcamp for data engineering for a lovely bunch of people, and we called it the Summer Camp. Here&apos;s &lt;a href=&quot;/2020/05/26/this-is-not-a-test-this-is-a-summer-camp-part-1.html&quot;&gt;part #1&lt;/a&gt; and &lt;a href=&quot;/2020/07/08/this-is-not-a-test-this-is-a-summer-camp-part-2.html&quot;&gt;part #2&lt;/a&gt; about how it all happened what we&apos;ve learned. (Btw: stay tuned, the dates for summer camp 2021 are already marked in our calendars…)&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;There are plenty of unspectacular things that kept us busy, however they were necessary: developing and testing the curriculum, dealing with Berlin&apos;s very own bureaucratic Kraken when founding a company, setting up corporate partnerships, updating Zoom every single day... just to name a few. The launch of our bootcamp was originally planned for the second half of 2020, but we&apos;ve decided to postpone as a result of the volatile outlook caused by corona. But spreading the word about what we&apos;re about to launch was never put on hold.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We&apos;ve started a &lt;a href=&quot;https://www.meetup.com/pipeline-data-engineering-academy-berlin/&quot;&gt;monthly Open House meetup session&lt;/a&gt; (virtually for now), so we can talk directly with folks who are interested in joining our training. It&apos;s happening every first Tuesday of the month. It helped us bridge the long cold Berlin winter without the beloved corona-compatible Data Stacks Picnics in the Park, which turned out to be a huge hit.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_stacks_picnic_in_the_park.jpg&quot; alt=&quot;People gathered around a table in a park&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;The need for clear guidance for navigating the world of the data engineering profession seems endless, so we put a lot of effort into sharing our POV. Daniel&apos;s monthly curated reading recommendations (aka &lt;a href=&quot;/the%20data%20janitor%20letters/2021/01/14/the-data-janitor-letters-december.html&quot;&gt;The Data Janitor Letters&lt;/a&gt;) received a lot of positive feedback from the data engineering community. Also, if you are trying to get access to the competences of a data engineer, check out &lt;a href=&quot;/2020/12/15/become-a-data-engineer-on-a-shoestring.html&quot;&gt;his highly successful guide about getting started with data engineering on a budget&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Yours truly has penned some words about &lt;a href=&quot;/2020/04/21/how-to-become-a-data-engineer.html&quot;&gt;learning paths&lt;/a&gt; for data engineers, about &lt;a href=&quot;/2020/09/03/the-state-of-coding-bootcamps-in-2020.html&quot;&gt;the state of bootcamps&lt;/a&gt; in general and about &lt;a href=&quot;/2020/09/22/data-engineer-salary-germany-2020.html&quot;&gt;data engineer salaries in Germany&lt;/a&gt; to increase transparency within the field.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Not enough, you say? We got you: &lt;a href=&quot;/2020/06/21/data-engineering-keynotes-to-remember.html&quot;&gt;here are some of Daniel&apos;s favourite keynotes&lt;/a&gt; and &lt;a href=&quot;/2020/08/07/data-engineering-data-janitor-keynotes.html&quot;&gt;you can see him on the stage of various conferences right here&lt;/a&gt;. Plus the chat with Tobias Macey from the &lt;a href=&quot;/2020/09/23/data-engineering-podcast-pipeline-academy.html&quot;&gt;Data Engineering Podcast&lt;/a&gt; about Pipeline Academy&apos;s pragmatic and commonsensical approach to data engineering challenges and the bootcamp itself.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Our merchandise was very 2020 as well:&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_training_course_learn_bootcamp_mask.jpg&quot; alt=&quot;A face mask printed with the Pipeline Academy logo&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Oh, I&apos;ve almost forgot: we&apos;ve launched &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;a short podcast series called #idataengineer&lt;/a&gt; to share the perspectives of other data engineers, and for Christmas we’ve opened the 24 tiny virtual windows of our &lt;a href=&quot;/2020/12/25/data-engineering-advent-calendar.html&quot;&gt;data engineering advent calendar&lt;/a&gt; with hints and drops of engineering wisdom for everyone.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Looking at my calendar, we&apos;ve actually met a lot of tech and data professionals both in-person and via video calls. To those of them reading this: your encouragement and input has been essential for our progress, so shout out to you, the humble and decent people of the tech/data/startup scene of Berlin!&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We&apos;ve been even featured on the popular and trashy @bestofkleinanzeigen instagram account. Everything happened so fast...&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_best_of_kleinanzeigen_course.jpg&quot; alt=&quot;Instagram post from Best of Kleinanzeigen showing a Pipeline Academy advert&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;We&apos;ve started &lt;a href=&quot;/corporate-training/&quot;&gt;supporting companies with their upskilling programs&lt;/a&gt; and consult data teams on data engineering topics like infrastructure design and architecture. It&apos;s an exciting journey and we&apos;re very happy to have smart and engaged startups and scale-ups in our growing clientele.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But our drive to explore and to learn goes even further: currently we&apos;re deeply involved in a transatlantic coaching-mentoring program, but this is something for the 2021 recap I guess...&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;At the very end of 2020 we&apos;ve been greenlit as an &lt;a href=&quot;/2021/01/10/bildungsgutschein-ready.html&quot;&gt;official partner of the Agentur für Arbeit&lt;/a&gt;, which means that we can support people who have been made redundant during the pandemic or have to work in Kurzarbeit. &lt;em&gt;Bildungsgutschein&lt;/em&gt; is the name of the game.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Kicking off 2021 on a high note:&lt;a href=&quot;/2021/01/10/sustainable-data-engineering.html&quot;&gt; incorporating sustainability into our methodology and data engineering curriculum&lt;/a&gt; is something we’re extremely proud of, and we truly believe that it will deliver tangible benefits to our future bootcamp participants.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;And there is so much more to come in 2021.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Stay healthy, folks.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_course_training_podcast.jpg&quot; alt=&quot;A podcast recording setup&quot;&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #9: Ayan Putatunda</title>
   <link href="https://dataengineering.academy/idataengineer/2021/01/26/idataengineer-confessions-interview-009.html"/>
   <updated>2021-01-26T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/01/26/idataengineer-confessions-interview-009.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/CZ5cE7AjsfM?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/9_ayan_wide_cover.jpg&quot; alt=&quot;Ayan, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - December 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2021/01/14/the-data-janitor-letters-december.html"/>
   <updated>2021-01-14T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2021/01/14/the-data-janitor-letters-december.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210121111745/https://www.hashpath.com/2020/12/an-analytics-engineer-is-really-just-a-pissed-off-data-analyst/&quot; target=&quot;_blank&quot;&gt;An analytics engineer is really just a pissed off data analyst&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Seth Rosen, Co-founder, Hashpath &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;... who has the tools and motivation to make things better for everyone else.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tkaszuba.medium.com/avro-schema-evolution-strategies-on-kafka-3c072a9a5347&quot; target=&quot;_blank&quot;&gt;Avro Schema Evolution Strategies on Kafka&lt;/a&gt;&lt;br&gt;&lt;em&gt;Tomasz Kaszuba, Java Big Data Engineer, Swiss Re&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I think the only place that the schema registry makes sense is when controlling 3rd party connections or in simple Kafka architectures.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://sudhir.io/the-big-little-guide-to-message-queues/&quot; target=&quot;_blank&quot;&gt;The Big Little Guide to Message Queues&lt;/a&gt;&lt;br&gt;&lt;em&gt;Sudhir Jonathan, Technical Architect, Qube Cinema&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Fundamental concepts that underlie them, and how they apply to popular queueing systems available today.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.dolthub.com/blog/2020-12-28-join-planning/&quot; target=&quot;_blank&quot;&gt;Planning joins to make use of indexes&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Zach Musgrave, Software Engineer, DoltHub&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Dolt is Git for Data. It&apos;s a SQL database that you can clone, fork, branch, and merge.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://evrl.com/devops/cloud/2020/12/18/serverless.html&quot; target=&quot;_blank&quot;&gt;Back to the &apos;70s with Serverless&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Cees de Groot, Principal Software Engineer, Canary Monitoring, Inc.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;History will repeat itself, of course.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210117091009/https://www.marketpulse.tech/tech/quotes-streamer&quot; target=&quot;_blank&quot;&gt;Going from 5K to 3M Messages/sec with 2ms Latency&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Market Pulse&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The evolution, failures and design decisions behind one of the world’s largest real-time, high-frequency and low-latency streaming systems.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20210115205742/https://blog.tomilkieway.com/72k-1/&quot; target=&quot;_blank&quot;&gt;We Burnt $72K testing Firebase + Cloud Run and almost went Bankrupt&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Sudeep Chauhan, Founder, Milkie Way Inc.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This is the story of how close we came to shutting down before even launching our first product, how we survived, and the lessons we learnt.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Introducing: Sustainable Data Engineering</title>
   <link href="https://dataengineering.academy/2021/01/10/sustainable-data-engineering.html"/>
   <updated>2021-01-10T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/01/10/sustainable-data-engineering.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Starting 2021 we&apos;re enhancing the educational focus of our data engineering bootcamp by integrating sustainability as a core value in addition to transparency, pragmatism and collaborative way of doing things.&lt;/p&gt;
&lt;blockquote&gt;&lt;h4 &gt;Pipeline Academy graduates will become the first data professionals equipped with practical data engineering concepts and best practices that support the environmental, economic and social sustainability of the products they build and the companies they join.&lt;/h4&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;We do this to empower our participants to use their engineering skills consciously taking all stakeholders of the ecosystem into consideration, but also to push data engineering culture in general to the next level. We&apos;re more than certain that this expertise will translate into a significant competitive advantage on the job market for our graduates.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But what does sustainability mean in the world of data engineering?&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/sustainability_data_engineering.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;1) Environmental sustainability in data engineering&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The architectures designed by data engineers leave a footprint on our planet: the most obvious but often ignored factor when setting up a data stack is energy consumption. Understanding the metrics and core drivers of the ecologic impact is essential for every environmentally conscious engineer. Just imagine, what if architects would ignore the quality and quantity of construction materials and processes they use for their buildings? What about managing data waste (hardware and software)? Have you ever &lt;a href=&quot;https://www.wired.com/story/amazon-google-microsoft-green-clouds-and-hyperscale-data-centers/&quot;&gt;compared how green the major cloud providers are compared to each other&lt;/a&gt;? Applying this know-how responsibly results in more efficient and sustainable data systems.&lt;/p&gt;
&lt;h4 &gt;2) Economic sustainability in data engineering&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Building and scaling infrastructure can quickly become a serious cost driver for organisations collecting and leveraging data actively. The blurry and deliberately incomparable pricing structures of competing data management tools make decision-makers without a clear guiding framework face a tricky situation with high financial risks. Applying the right KPIs, anticipating business needs and considering technical constraints can lead to meaningful cost reductions for tooling, and the contribution of the responsible experts will support the long-term success of their businesses. Our advisor, Dr. Martin Loetzsch is the Chief Data Officer at Project A Ventures, and their &lt;a href=&quot;https://www.sustainability-playbooks.com/&quot;&gt;Sustainability Playbooks&lt;/a&gt; are among the useful guiding lights for us.&lt;/p&gt;
&lt;h4 &gt;3) Social sustainability in data engineering&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;For the last decade, data has been treated as the new oil for economic growth. But just like with fossil fuel, there are certain social consequences of managing and using data for business purposes. Even in the current regulatory environment (that is expected to become more and more strict in the near future), data engineers have one of the most central roles when it comes to informing stakeholders and executing engineering work according to guidelines like GDPR (General Data Protection Regulation) and CCPA (California Consumer Privacy Act). Addressing privacy and security concerns proactively, understanding the potential for abuse when working with user data, having a data management/governance in place are just a few of the concerns data professionals have to deal with on a daily basis. But it does not stop there: creating a healthy working environment for data teams by managing roles and responsibilities, increasing the transparency for their work to manage expectations are just as important for sustainable careers and successful work. To name an example, Maciej Cegłowski of Pinboard &lt;a href=&quot;https://idlewords.com/talks/haunted_by_data.htm&quot;&gt;talked about this back in 2015&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Why are we doing this?&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/sustainable_data_engineering.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Your opportunity&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;The primary purpose of Pipeline Academy is to help individuals learn &lt;a href=&quot;/curriculum/&quot;&gt;data engineering fundamentals&lt;/a&gt; so they can improve on their career opportunities. Daniel and I focus a lot on ideas that increase the long-term effectiveness of the content and methodologies we use. When observing the tectonic shift in consumer behaviour towards a more environmentally and socially conscious economy, we can see that it presents a great opportunity for every professional today. Learning how to create products or manage a business so they do little to no harm to others is becoming a huge competitive advantage on the job market, and will eventually be essential for tomorrow&apos;s purpose-driven leaders. In simple terms: we give our bootcamp participants not only the technical know-how about building data infrastructure, but also context about their responsibilities and options for managing its impact.&lt;/p&gt;
&lt;h3 &gt;Our shared responsibility&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Data engineers set up infrastructure that impact our social and ecological environment, and it is their duty to design systems that consider the ecosystem of stakeholders it is interacting with. Designing software architectures and optimising operational efforts will become much more pronounced mid term: employers selling i.e. eco-friendly products are going to hire data professionals who can execute on this specific value proposition. As educators, we have the power to equip the future architects and builders of data platforms with the knowledge that enables them to create a better working and living environment. They will build the systems that power the technological reforms in mobility, healthcare and education in the upcoming years, and they will educate future generations of data experts - and this multiplier effect makes us even more aware of our influence. &lt;a href=&quot;https://2019.stateofeuropeantech.com/&quot;&gt;Check how your company or a company you&apos;re planning to join does in these terms&lt;/a&gt;.&lt;/p&gt;
&lt;h3 &gt;Rooted in our resourcefulness&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;In the last 15 years hyper-growth has become the standard expectation towards tech companies, which resulted in an unsustainable business- and working culture that does not support the long-term coexistence of social and environmental stakeholders. With Pipeline Academy we&apos;re building a bootstrapped educational institution that we&apos;d like to establish long-term, and we won&apos;t sacrifice our lean DIY-attitude and minimalistic approach for growth. Conscious environmental management of our own organisation, symbiotical coexistence with our partners, and actively fostering a sustainable data and learning culture with our educational concept are proof of that. Compared to some other programs we put the emphasis on long-lasting value in handmade quality with a focus on the individual, and avoid ephemeral and trend-driven promises addressing the mass market.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/sustainable_data.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Sustainable data engineering in practice&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;The sustainability concept is integrated into our curriculum in the form of weekly lectures, discussions, case studies, reading materials, but also very hands-on exercises. Some of the topics and concepts we discuss in the bootcamp:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Designing lean data architectures - explain it to your grandma,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Unix philosophy and the Zen of Python,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;KISS, figure out what to worry about, learn what not to do,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&apos;Take the long view&apos; (Prof. Galloway), invest in durable knowledge,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Strive for fast and good enough (Pareto and YAGNI),&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Code longevity, reuse and recycle code,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;System optimisation: remove, retire and delete,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Measuring tooling efficiency: speed, data storage and energy consumption,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Just because you can it does not mean you should - you are not Google, you don&apos;t need their stack,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Don&apos;t feed the FAANG beast with your data.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;This is your chance to get ahead of the curve and learn data engineering enhanced with principles that contribute to a more sustainable software development and data management culture.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;PS: We too are learning and exploring along the way, so if you have suggestions for us on how to improve on the subject matter, or would like to cooperate - do let us know! We&apos;re always open for meaningful collaborations.&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Bildungsgutschein-ready</title>
   <link href="https://dataengineering.academy/2021/01/10/bildungsgutschein-ready.html"/>
   <updated>2021-01-10T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2021/01/10/bildungsgutschein-ready.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Pipeline Academy has received the coveted AZAV-certification, so from now on you can join our bootcamp with a Bildungsgutschein issued by the Agentur für Arbeit (AfA)! This means basically that in case you are between jobs, have been put on Kurzarbeit or are at risk of losing your current employment, you are very likely to be eligible for a voucher (the Bildungsgutschein) from the German AfA, so you can join our bootcamp &lt;em&gt;free of charge&lt;/em&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Since the launch of Pipeline Academy, our mantra was making data engineering more accessible. This is indeed a huge step both for us towards that direction, just as for the people who would make a great fit for our course but could not afford taking part until now. This is our guide to help you leverage this life-changing opportunity and putting yourself on a sustainable career path.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The Bildungsgutschein can cover your full tuition fee if:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You are unemployed receiving ALG I.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You are on Kurzarbeit&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;You are employed and are eligible for QCG (&lt;a href=&quot;https://www.bgbl.de/xaver/bgbl/start.xav?startbk=Bundesanzeiger_BGBl&amp;amp;jumpTo=bgbl118s2651.pdf#__bgbl__%2F%2F*%5B%40attr_id%3D%27bgbl118s2651.pdf%27%5D__1609684990242&quot;&gt;Qualifizierungschancengesetz&lt;/a&gt;)&lt;br&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;Getting a Bildungsgutschein when unemployed&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The &lt;a href=&quot;https://www.arbeitsagentur.de/Bildungsgutschein&quot;&gt;Bildungsgutschein&lt;/a&gt; is a form of direct financial support provided by the German federal state for those individuals who seek support for acquiring new skills at certified educational institution. The goal of this instrument is to help people avoid unemployment or to enable individuals to (re-)enter to the job market successfully and in a sustainable way. It&apos;s used either for broadening an existing skillset or for switching careers and starting a journey in a new profession.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The process of receiving a Bildungsgutschein is fairly straightforward: check if you are eligible, select a training program matching your goals, present your case and enroll with your voucher. This is a so called &quot;Soll-Leistung&quot;, which translates into an optional offering: based on an evaluation of your individual situation by your designated clerk at the AfA, it will be determined whether you can receive a Bildungsgutschein for the course of your choice. However, the current economic situation forces a lot of people to find a new occupation and the Jobcenter recognizes that by supporting them accordingly.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;The process looks something like this:&lt;/strong&gt;&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Schedule a call with us to discuss your plan of joining Pipeline Academy, we&apos;re happy to give you advice on how to obtain a Bildungsgutschein.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Register with the Agentur für Arbeit or Jobcenter (if you haven&apos;t done so yet) and schedule a consulting session with your contact person to discuss your eligibility and your plan.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p class=&quot;&quot; &gt;Prepare your case for choosing data engineering as your desired future career path and Pipeline Academy, and be ready to answer questions like:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What does a data engineer do?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What convinced you to become a data engineer?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Why is Pipeline Academy the best school to learn data engineering?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;&lt;li&gt;&lt;p class=&quot;&quot; &gt;What kind of positions are you going to apply for after the bootcamp (e.g. data engineer, BI engineer, machine learning engineer, data platform manager)?&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;The market demand for data engineering skills and the number of open positions are very convincing arguments for AfA advisors, as they want to see their clients succeed by getting a great a job.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Once you&apos;ve received the green light from the AfA, let us know so we can secure your spot in one of our upcoming cohorts and we&apos;ll also take care of the necessary paperwork.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Clear your schedule for the 12 weeks of the bootcamp and prepare for learning something that is definitely going to expand your professional horizons!&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h4 &gt;A couple of relevant points to keep in mind&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Do your research&lt;/strong&gt;: take the time to understand the role of a data engineer, look at the profiles of companies hiring and the trends behind the growing demand for data engineering competences. If you have other training options you are considering, make sure to talk to them to have a better overview and get a feeling for how it feels to be enrolled in their classes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Prepare for the interview with the AfA&lt;/strong&gt;: the goal of the AfA is to secure your employment for the upcoming years. If you have figured what path fits to you and what you would like to do in the future professionally, make sure to be able to explain your plan in a simple and clear manner.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Commit&lt;/strong&gt;: the opportunity that is presented to you via a free Bildungsgutschein is something that you should not take lightheartedly. We&apos;re here to do everything in our power to help you get started in data engineering, but we expect you to show up and put in the work: this is how you can set yourself up for a new career.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/bildungsgutschein_data_engineering_bootcamp_berlin.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;If you are on Kurzarbeit&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;During the COVID-19 lockdown a lot of businesses were forced to make difficult decisions in order to secure their existence. One of the measures that has proven very effective was the introduction of &lt;a href=&quot;https://www.bmas.de/SharedDocs/Downloads/DE/kug-faq-kurzarbeit-und-qualifizierung-englisch.pdf?__blob=publicationFile&amp;amp;v=5&quot;&gt;Kurzarbeit&lt;/a&gt;: employers who suffered a hit during the economic crisis have the opportunity to secure financial compensation for their employees from the federal state in exchange for limiting their working hours. It means practically that employees get to keep their jobs, work less hours, but get reimbursed similarly to as if they would work full-time - this is how massive short-term layoffs are avoided. If you find yourself in such a situation, you might have the time and the chance to learn something new with the support of the Agentur für Arbeit in form of a Bildungsgutschein. Make sure to check-in with your employer&apos;s HR, they will support you in the process, but the steps to get there are pretty much the same as outlined above.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;QCG aka Qualifizierungschancengesetz&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This measure creates a unique win-win-win situation. It&apos;s &lt;a href=&quot;https://www.arbeitsagentur.de/unternehmen/finanziell/foerderung-von-weiterbildung&quot;&gt;aimed to help employees at risk of losing their jobs&lt;/a&gt;: in cooperation with their employers they can participate in a bootcamp that makes it possible for them to avoid unemployment on the long run. On top of funding the training for the employee, &lt;a href=&quot;https://www.arbeitsagentur.de/m/weiterbildung-qualifizierungsoffensive/&quot;&gt;the employer receives financial support as well&lt;/a&gt; to cover the salary of the employee and social security for during the training. For this, we encourage employees to approach the responsible departments in their organisations to start the application process with the AfA. We also support employers looking for meaningful upskilling for their staff.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/bildungsgutschein_coding_bootcamp.jpg&quot; alt=&quot;Two people working through something together at a laptop&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Make sure to look into all of the available &lt;a href=&quot;/financing-scholarships/&quot;&gt;financing options&lt;/a&gt; we provide so you can make an educated decision about this realm. Depending on your eligibility, your current financial situation and future plans you should pick the financing method that fits best to your needs. If you&apos;d like to learn more or have a chat about the pros and cons, feel free to &lt;a href=&quot;/apply/&quot;&gt;contact us or schedule a call&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #8: Oleg Soroka</title>
   <link href="https://dataengineering.academy/idataengineer/2021/01/07/idataengineer-confessions-interview-008.html"/>
   <updated>2021-01-07T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2021/01/07/idataengineer-confessions-interview-008.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/5lkxHPVD-lw?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;div class=&quot;sqs-video-overlay&quot; style=&quot;opacity: 0;&quot;&gt;&lt;img src=&quot;/images/posts/8_oleg_wide_cover-copy.jpg&quot; alt=&quot;Oleg, Pipeline Academy participant&quot;&gt;&lt;div class=&quot;sqs-video-opaque&quot;&gt;&lt;/div&gt;&lt;div class=&quot;sqs-video-icon&quot;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #7: Adrian Brudaru</title>
   <link href="https://dataengineering.academy/idataengineer/2020/12/30/idataengineer-confessions-interview-007.html"/>
   <updated>2020-12-30T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/12/30/idataengineer-confessions-interview-007.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/QydCoMRy6Jc?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Advent Calendar 2020</title>
   <link href="https://dataengineering.academy/2020/12/25/data-engineering-advent-calendar.html"/>
   <updated>2020-12-25T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/12/25/data-engineering-advent-calendar.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Throughout December 2020 we’ve shared a daily dose of semi-esoteric data engineering wisdom on our social media channels (&lt;a href=&quot;https://www.instagram.com/dataengineering.academy/&quot;&gt;instagram&lt;/a&gt; and &lt;a href=&quot;https://www.linkedin.com/school/pipeline-data-engineering-academy/&quot;&gt;LinkedIn&lt;/a&gt;). This post shall serve as a commemorative monolith you can always turn to when the data engineering gods are not picking up your call.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201201.png&quot; alt=&quot;Data engineering advent calendar, day 1&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#1: Don’t write code, solve the problem.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201202.png&quot; alt=&quot;Data engineering advent calendar, day 2&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#2: Python is always at hand to pretty print a JSON:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ python3 -m json.tool some.json&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201203.png&quot; alt=&quot;Data engineering advent calendar, day 3&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#3: &lt;em&gt;EXPLAIN&lt;/em&gt; is your friend.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201204.png&quot; alt=&quot;Data engineering advent calendar, day 4&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#4: &quot;Choose boring technology.&quot; Dan McKinley @mcfunley&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201205.png&quot; alt=&quot;Data engineering advent calendar, day 5&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#5: Complicated is better than complex.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201206.png&quot; alt=&quot;Data engineering advent calendar, day 6&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#6: What do you do on your CLI?&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ &amp;lt; ~/.bash_history | sort | uniq -c | sort -n&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201207.png&quot; alt=&quot;Data engineering advent calendar, day 7&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#7: SQLite (2000) has one trillion (1e12) active installs. It&apos;s a file with SQL API and window functions.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201208.png&quot; alt=&quot;Data engineering advent calendar, day 8&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#8: &quot;Premature optimization is the root of all evil.&quot; Tony Hoare (Quicksort)&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201209.png&quot; alt=&quot;Data engineering advent calendar, day 9&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#9: Keep It Simple (&amp;amp;) Stupid and remember separation of concerns.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201210.png&quot; alt=&quot;Data engineering advent calendar, day 10&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#10: Log in to a recently launched container (via @basmatitree)&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ docker exec -it $(docker ps -q | tail -1) /bin/bash&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201211.png&quot; alt=&quot;Data engineering advent calendar, day 11&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#11: Get details of last failed Redshift load.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;SELECT * FROM stl_load_errors ORDER BY starttime DESC LIMIT 1;&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201212.png&quot; alt=&quot;Data engineering advent calendar, day 12&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#12: &quot;When in doubt, use brute force.&quot; Ken Thompson (Go, UTF-8, Unix)&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201213.png&quot; alt=&quot;Data engineering advent calendar, day 13&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#13: Maintainable is debuggable and testable, and has version control.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201214.png&quot; alt=&quot;Data engineering advent calendar, day 14&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#14: Remove all local git branches other than master and the currently used one (via @advincze)&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ git branch --no-color | grep -v &apos;master&apos; | grep -v &apos;*&apos; | xargs git branch -D&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201215.png&quot; alt=&quot;Data engineering advent calendar, day 15&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#15: Order of execution in SQL:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;FROM WHERE GROUP BY HAVING SELECT [DISTINCT] UNION ORDER BY&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201216.png&quot; alt=&quot;Data engineering advent calendar, day 16&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#16: &quot;Don&apos;t reinvent the flat tire.&quot; Alan Kay (Squeak, Smalltalk, OOP, GUI)&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201217.png&quot; alt=&quot;Data engineering advent calendar, day 17&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#17: Code is dependency. Others&apos; code is dependency squared. Delete, remove, retire.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201218.png&quot; alt=&quot;Data engineering advent calendar, day 18&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#18: You can run SQL directly on your CLI on CSV or TSV files with &lt;a href=&quot;http://harelba.github.io/q/&quot; target=&quot;_blank&quot;&gt;http://harelba.github.io/q/&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201219.png&quot; alt=&quot;Data engineering advent calendar, day 19&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#19: Queries on MPPs? Use WITH/CTEs, filter with WHERE, SELECT explicitly, avoid JOIN, SORTKEYS are your friends. It&apos;s all about scanning less.&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201220.png&quot; alt=&quot;Data engineering advent calendar, day 20&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#20: “Bad programmers worry about the code. Good programmers worry about data structures and their relationships.” Linus Torvalds (Git, Linux)&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201221.png&quot; alt=&quot;Data engineering advent calendar, day 21&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#21: The longer a technology lives, the longer it can be expected to live (Lindy effect)&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201222.png&quot; alt=&quot;Data engineering advent calendar, day 22&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#22: Delete files recursively:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ find . -name &quot;*.pdf&quot; -print0 | xargs -0 rm&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;

                  &lt;img src=&quot;/images/posts/20201223.png&quot; alt=&quot;Data engineering advent calendar, day 23&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#23: Why PostgreSQL (1996)? It is &lt;em&gt;the&lt;/em&gt; open source RDBMS with columnar (cstore), geo (PostGIS), timeseries (TimescaleDB) and REST API (PostgREST).&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/20201224.png&quot; alt=&quot;Data engineering advent calendar, day 24&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;#24: &quot;Use simple algorithms as well as simple data structures.&quot; Rob Pike (Go, UTF-8, Unix)&lt;/p&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - November 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/12/24/the-data-janitor-letters-november-2020.html"/>
   <updated>2020-12-24T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/12/24/the-data-janitor-letters-november-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.jacquescorbytuech.com/writing/marketers-addicted-bad-data&quot; target=&quot;_blank&quot;&gt;Marketers are Addicted to Bad Data&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Jacques Corby-Tuech, Marketing Operations Manager, CyberSmart &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;36% percent of people in the UK use an adblocker.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://simonwillison.net/2020/Nov/14/personal-data-warehouses/&quot; target=&quot;_blank&quot;&gt;Personal Data Warehouses: Reclaiming Your Data&lt;/a&gt;&lt;br&gt;&lt;em&gt;Simon Willison, Datasette&lt;/em&gt;
&lt;/h4&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/fandom-engineering/aws-s3-disaster-recovery-using-versioning-and-objects-metadata-6b9f56ac8858&quot; target=&quot;_blank&quot;&gt;AWS S3 — Disaster recovery using versioning and objects metadata&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jacek Małyszko, Data Engineer, Fandom&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Accidental removal of data on S3 is something that no Data Engineer on AWS wants to be involved in.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://brandonharris.io/redshift-clickhouse-time-series/&quot; target=&quot;_blank&quot;&gt;ClickHouse, Redshift and 2.5 Billion Rows of Time Series Data&lt;/a&gt;&lt;br&gt;&lt;em&gt;Brandon Harris, Cloud + Analytics, Discover Financial&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In this post I show you how to synthesize billions of rows of true time series data with an autoregressive component, and then explore it with ClickHouse, a big data scale OLAP RDBMS, all on AWS.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.datasyndrome.com/python-and-parquet-performance-e71da65269ce&quot; target=&quot;_blank&quot;&gt;Python and Parquet performance optimization using Pandas, PySpark, PyArrow, Dask, fastparquet and AWS S3&lt;/a&gt;&lt;br&gt;&lt;em&gt;Russell Jurney, Principal Consultant, Data Syndrome&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This post outlines how to use all common Python libraries to read and write Parquet format while taking advantage of columnar storage, columnar compression and data partitioning.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.cloudflare.com/clickhouse-capacity-estimation-framework/&quot; target=&quot;_blank&quot;&gt;ClickHouse Capacity Estimation Framework&lt;/a&gt;&lt;br&gt;&lt;em&gt;Oxana Kharitonova, SRE, Cloudflare&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Our current insertion rate is about 90M rows per second.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://shopify.engineering/build-production-grade-workflow-sql-modelling&quot; target=&quot;_blank&quot;&gt;How to Build a Production Grade Workflow with SQL Modelling&lt;/a&gt;&lt;br&gt;&lt;em&gt;Michelle Ark, Senior Data Engineer, Shopify&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Currently, we have a warehouse consisting of over 100 models, and this validation step takes about two minutes.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://towardsdatascience.com/the-state-of-open-source-data-integration-and-etl-d2f2e8733e2a&quot; target=&quot;_blank&quot;&gt;The State of Open-Source Data Integration and ETL&lt;/a&gt;&lt;br&gt;&lt;em&gt;John Lafleur, Co-Founder, Airbyte&lt;/em&gt; &lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Is an open-source (OSS) approach is more relevant than a commercial software approach in addressing the data integration problem.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #6: Bence Faludi</title>
   <link href="https://dataengineering.academy/idataengineer/2020/12/23/idataengineer-confessions-interview-006.html"/>
   <updated>2020-12-23T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/12/23/idataengineer-confessions-interview-006.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/zkWuWWolvE4?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #5: Leandro Loi</title>
   <link href="https://dataengineering.academy/idataengineer/2020/12/16/idataengineer-confessions-interview-005.html"/>
   <updated>2020-12-16T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/12/16/idataengineer-confessions-interview-005.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/7K0dnxUC-2k?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Become a Data Engineer on a Shoestring (aka The Best Free Courses and Learning Resources)</title>
   <link href="https://dataengineering.academy/2020/12/15/become-a-data-engineer-on-a-shoestring.html"/>
   <updated>2020-12-15T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/12/15/become-a-data-engineer-on-a-shoestring.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;I was tinkering with the idea of finding the right way to help others to identify the resources that give you bang for the buck when it comes to upskilling yourself in data engineering… and this is what I came up with. So this is how to spend the remainder of your learning budget before the year ends.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.ted.com/talks/alan_kay_a_powerful_idea_about_ideas/up-next&quot;&gt;Alan Kay&lt;/a&gt; (Squeak, Smalltalk, OOP, GUI) suggested that programming is &lt;a href=&quot;https://lispcast.com/programming-pop-culture/&quot;&gt;pop culture&lt;/a&gt;, because it spreads much, much faster than mentorship, education or formal study does. I think the best example of this are the &lt;a href=&quot;http://soobrosa.info/static/2016-04-05_my-humble-james-mickens-shrine-aka-the-only-real-combined-cs-degree-and-mba-you-will-ever-need.html&quot;&gt;talks of James Mickens&lt;/a&gt;, they can help shine a light on why people think that blockchain or machine learning are &lt;em&gt;the&lt;/em&gt; solution.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Treat this as a restaurant menu. No one should consume it end-to-end. Scan it and pick something that looks good. Try it. Rinse and repeat! Wishing you a great 2021!&lt;/p&gt;
&lt;h4 &gt;tl;dr&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;A) I want to pass the interview for ...&lt;/strong&gt; &lt;br&gt;Learn Spark or Kafka or whatever people do these days at a &lt;a href=&quot;https://www.coursera.org/courses?languages=en&amp;amp;query=data%20science%20specialization&quot;&gt;Coursera Specialization&lt;/a&gt; or a &lt;a href=&quot;https://www.udacity.com/school-of-data-science&quot;&gt;Udacity Nanodegree&lt;/a&gt; — I did them, I did not like them. Expect to burn about a hundred bucks.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;B) I want a band aid aka branded lock-in trainings!&lt;/strong&gt;&lt;br&gt;Cloud providers, and OS-as-a-marketing-tool companies like &lt;a href=&quot;https://www.cloudera.com/about/training.html#find-training&quot;&gt;Cloudera&lt;/a&gt; or &lt;a href=&quot;https://www.databricks.com/learn/training/home&quot;&gt;Databricks&lt;/a&gt; are there to help you! Either for free or for a thousand dollars.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;C) Gamify my learning!&lt;/strong&gt;&lt;br&gt;For browser-based self-paced bytesize Sudoku, try learning Python and SQL in a command line at &lt;a href=&quot;https://www.dataquest.io/&quot;&gt;DataQuest&lt;/a&gt; or at &lt;a href=&quot;https://www.datacamp.com/&quot;&gt;DataCamp&lt;/a&gt; — the latter one is also available on &lt;a href=&quot;https://www.datacamp.com/mobile/&quot;&gt;mobile&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;D) I have to be able to bluff my way into mastering data engineering tools by tomorrow!&lt;br&gt;&lt;/strong&gt;&lt;a href=&quot;https://www.udemy.com/&quot;&gt;Udemy&lt;/a&gt; is the &apos;you get what you pay for&apos; eBay kleinanzeigen for your needs. 10 bucks will cover most courses when they are on sale == that&apos;s basically always.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_books.jpg&quot; alt=&quot;A shelf of programming books&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Learn SQL&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;It is back.&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The single most important technology that anybody who has &apos;data&apos; in her title must master.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://selectstarsql.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;Select Star SQL&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (free, 3 hours)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;An interactive book for learning SQL with SQLite running &lt;a href=&quot;https://github.com/kripken/sql.js&quot;&gt;&lt;em&gt;in the browser&lt;/em&gt;&lt;/a&gt; - no setup required, the best intro for a beginner.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://mystery.knightlab.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;The SQL Murder Mystery&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (free)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;A good way to practice SQL skills after the tutorial, works in the browser.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://sqlpd.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;SQL Police Department&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; ($20)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;If you got a taste for story-based learning, continue here.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;Learn Programming&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The uncrowned coding intro is definitely the&lt;strong&gt;&lt;em&gt; &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://www.edx.org/course/cs50s-introduction-to-computer-science&quot;&gt;&lt;strong&gt;&lt;em&gt;Harvard edX CS50&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (free or €167 for a Verified Certificate)&lt;/em&gt;&lt;/strong&gt; or you can learn at &lt;strong&gt;&lt;em&gt;Dr. Chuck&apos;s &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://www.coursera.org/specializations/python&quot;&gt;&lt;strong&gt;&lt;em&gt;Python for Everybody&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; at Coursera&lt;/em&gt;&lt;/strong&gt; (you can do it in a focused month for €41).&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;If you have previous programming exposure you could spend 4 days on honing your Python skills with&lt;strong&gt;&lt;em&gt; &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://github.com/dabeaz-course/practical-python&quot;&gt;&lt;strong&gt;&lt;em&gt;Practical Python Programming&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; by David Beazley&lt;/em&gt;&lt;/strong&gt;. I totally recommend his presentations/workshops&lt;strong&gt;&lt;em&gt; &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://web.archive.org/web/20201109032141/http://www.dabeaz.com/generators/Generators.pdf&quot;&gt;&lt;strong&gt;&lt;em&gt;Generator Tricks For Systems Programmers&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; and &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://web.archive.org/web/20201214095936/http://www.dabeaz.com/coroutines/Coroutines.pdf&quot;&gt;&lt;strong&gt;&lt;em&gt;A Curious Course on Coroutines and Concurrency&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. I was about to embed his Discovering Python keynote, but I&apos;d better just point to that it&apos;s one of my favourite &lt;a href=&quot;/2020/06/21/data-engineering-keynotes-to-remember.html&quot;&gt;keynotes&lt;/a&gt; ever.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;Watch screencasts&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I didn’t understand the format until Torsten showed me Destroy All Software. It&apos;s not for everyone, it&apos;s not for everything, but now I understand its place in the universe. And yes, you can even learn computer science with it. Computer science is actually fun, too bad you don&apos;t get to do it much in the everyday crunch — although it pays well if you know its most crucial concepts.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.destroyallsoftware.com/screencasts&quot;&gt;&lt;strong&gt;&lt;em&gt;Destroy All Software&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; ($29/month)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Classic is mostly Ruby, computation is Python. Can&apos;t miss the Unix and Git ones.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.executeprogram.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;execute program&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (16 free lessons then $19/month)&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;Hit the SQL and the Regex.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;&lt;em&gt;Computer Science Essentials&lt;/em&gt;&lt;/strong&gt;&lt;strong&gt;&lt;em&gt;, (&lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://web.archive.org/web/20210128104427/https://bigmachine.io/free/&quot;&gt;&lt;em&gt;free samples&lt;/em&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;, $79)&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;You need Database Normalization, Make, Reading Shell Scripts, Shell Script Basics. Has a book version, see below at &lt;em&gt;The Imposter&apos;s Handbook&lt;/em&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://missing.csail.mit.edu/&quot;&gt;&lt;strong&gt;&lt;em&gt;The Missing Semester of Your CS Education&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (free)&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;Check Shell Tools and Scripting, Data Wrangling, Command-line Environment, Version Control (Git), Debugging and Profiling.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_books-nn1741.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Top 3 books I recommend&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.datascienceatthecommandline.com/2e/&quot;&gt;&lt;strong&gt;&lt;em&gt;Jeroen Janssens: Data Science at the Command Line, 2e&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (O’Reilly)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Not just data science at the command line. The new one with &lt;code&gt;make&lt;/code&gt; will take you very, very far.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://pragprog.com/titles/bksqla/sql-antipatterns/&quot;&gt;&lt;strong&gt;&lt;em&gt;Bill Karwin: SQL Antipatterns - Avoiding the Pitfalls of Database Programming&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (Pragmatic Bookshelf)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;When I got to the third chapter describing situations that I encountered personally in my career I was 100% sold. Also for all who think database programming is not a thing. Hold my beer for a sec.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;&lt;em&gt;Greg Wilson: Data Crunching, 2005 (Pragmatic Bookshelf)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;My first book on data. No frills, talks about file formats, has code examples, hits regex. Every day you will do something that is covered in this book.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;More books for the data engineer&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://bigmachine.io/products/imposter-season-1/&quot;&gt;&lt;strong&gt;&lt;em&gt;The Imposter&apos;s Handbook&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (each book $30)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;You need the first book: Database Normalization, Make, Reading Shell Scripts, Shell Script Basics. The second one is a good-to-have. See &lt;em&gt;Computer Science Essentials&lt;/em&gt; above for the video version.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://bigmachine.io/products/a-curious-moon/&quot;&gt;&lt;strong&gt;&lt;em&gt;A Curious Moon&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; ($30)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Beautiful ebook with SQL exercises, also intro to Postgres with real NASA data.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;&lt;em&gt;Markus Winand: SQL Performance Explained (&lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://sql-performance-explained.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;book&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; for money, &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://use-the-index-luke.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;free web version&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;A multiformat, practical handbook on advanced SQL and running databases in general.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;A free &lt;a href=&quot;https://greenteapress.com/wp/think-python-2e/&quot;&gt;Allen B. Downey: Think Python&lt;/a&gt; 2nd Edition and a $6.99 &lt;a href=&quot;https://leanpub.com/learnbashthehardway&quot;&gt;Ian Miell: Learn Bash the Hard Way&lt;/a&gt; ebook wouldn&apos;t hurt to have either.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;Books for the connoisseur&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;&lt;em&gt;Julia Evans: &lt;/em&gt;&lt;/strong&gt;&lt;a href=&quot;https://wizardzines.com/&quot;&gt;&lt;strong&gt;&lt;em&gt;programming zines&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Important topics are covered, and it will show how you can talk about concepts without fuzz. I like them all. Check the black and white free ones.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.oreilly.com/library/view/data-analysis-with/9781449389802/&quot;&gt;&lt;strong&gt;&lt;em&gt;Philipp K. Janert: Data Analysis with Open Source Tools&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (O’Reilly)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;High-end analysis practice, more useful than most ML/AI. Very well written. It raised the bar for me in technical writing.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.amazon.de/-/en/Angus-Croll/dp/1593275854&quot;&gt;&lt;strong&gt;&lt;em&gt;Angus Croll: If Hemingway Wrote JavaScript&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt;&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;Only if you&apos;re into languages and the thinking structures they impose, but then you will start gifting it.&lt;br&gt;&lt;/p&gt;
&lt;h4 &gt;Books for your Zoom background&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://dataintensive.net/&quot;&gt;&lt;strong&gt;&lt;em&gt;Martin Kleppmann: Designing Data-Intensive Applications&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; (O’Reilly)&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;I wish this would be our problem, not &lt;a href=&quot;https://books.google.de/books?id=zFheDgAAQBAJ&amp;amp;pg=PR5&amp;amp;dq=alan+kay+pop+culture&amp;amp;hl=en&amp;amp;sa=X&amp;amp;ved=2ahUKEwi2tIWV14ntAhWMecAKHTubBs0Q6AEwA3oECAEQAg#v=onepage&amp;amp;q=alan%20kay%20pop%20culture&amp;amp;f=false&quot;&gt;pop culture&lt;/a&gt;. Great and detailed, limited relevance from a practical point of view. Very academic, you can watch ‘em &lt;a href=&quot;https://www.youtube.com/playlist?list=PLeKd45zvjcDFUEv_ohr_HdUFe97RItdiB&quot;&gt;videos&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.kimballgroup.com/data-warehouse-business-intelligence-resources/books/data-warehouse-dw-toolkit/&quot;&gt;&lt;strong&gt;&lt;em&gt;Ralph Kimball and Margy Ross: The Data Warehouse Toolkit&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; 3rd Edition &lt;/em&gt;&lt;/strong&gt;&lt;br&gt;I mean it&apos;s too long, it&apos;s too slow, it belongs to a previous era, where computing has been immensely less performant, still if you want to build mental models of how data describes different industries, a good old printed SAP.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.dabeaz.com/cookbook.html&quot;&gt;&lt;strong&gt;&lt;em&gt;Python Cookbook&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;&lt;em&gt; 3rd Edition&lt;br&gt;&lt;/em&gt;&lt;/strong&gt;One has to have a reference book on Python. You know, no Internet, no Stack Overflow.&lt;br&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Thanks to &lt;a href=&quot;https://twitter.com/t0rsten&quot;&gt;Torsten Becker&lt;/a&gt;, &lt;a href=&quot;/about/&quot;&gt;Péter Fábián&lt;/a&gt;, &lt;a href=&quot;https://www.liberties.eu/en/info/our-team/12&quot;&gt;Balázs Dénes&lt;/a&gt; and &lt;a href=&quot;https://twitter.com/skiedel&quot;&gt;Steffen Kiedel&lt;/a&gt; for discussions over lunches in Berlin.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Support gastronomy during lockdown, order food from your valued vendors.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;One more thing:&lt;br&gt;Check out the &lt;/em&gt;&lt;a href=&quot;/2022/02/18/learn-data-engineering-on-a-shoestring-free-courses.html&quot;&gt;&lt;em&gt;second edition of this series&lt;/em&gt;&lt;/a&gt;&lt;em&gt; with even more recommended learning experiences for aspiring data engineers!&lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_books-4frzr1.jpg&quot; alt=&quot;&quot;&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #4: Zoltán C. Tóth</title>
   <link href="https://dataengineering.academy/idataengineer/2020/12/09/idataengineer-confessions-interview-004.html"/>
   <updated>2020-12-09T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/12/09/idataengineer-confessions-interview-004.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it. &lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/gVidOG6wqD0?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #3: Sejal Vaidya</title>
   <link href="https://dataengineering.academy/idataengineer/2020/12/03/idataengineer-confessions-interview-003.html"/>
   <updated>2020-12-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/12/03/idataengineer-confessions-interview-003.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/GELFXfcIJH4?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>A data prediction for 2021</title>
   <link href="https://dataengineering.academy/2020/11/26/data-trend-report-prediction-for-2021.html"/>
   <updated>2020-11-26T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/11/26/data-trend-report-prediction-for-2021.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Yes, it&apos;s &lt;em&gt;that&lt;/em&gt; time of the year again: in the coming weeks ultra-scientific research institutes, cutting-edge management and innovation consultancies and basically everyone with a podcast are going to share their predictions about the most important technology trends for 2021. In a parallel universe called the real world, businesses have to make their own predictions for next year as well (for revenues, staffing, organizational changes etc.), they just prefer to call it &lt;em&gt;planning&lt;/em&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Thinking about scenarios for what&apos;s ahead carries tremendous value, this has been known for long in the military even &lt;a href=&quot;https://quoteinvestigator.com/2017/11/18/planning/&quot;&gt;before Dwight D. Eisenhower&lt;/a&gt; famously said:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Plans are worthless, but planning is everything”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s about preparing for what is likely to happen and somehow for the unknowns that might impact our vague idea of how the future will unfold. The value I&apos;ve mentioned is in part generated by the fact that anticipating certain future events today and acting in a controlled way so that these events result in positive (or not-negative) outcomes make a business more likely to survive the stormy seas of an economy shaken by the pandemic. But there is also value in the process itself, as it&apos;s an exercise in preparation.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;So while these lists are being released in the upcoming weeks, the name of the game for the readers is finding patterns that help uncover the challenges on the radar and navigating through the various agendas inherently built into them. By doing this, one might arrive at something that&apos;s fairly likely to end up a real trend in the near future, not just an empty buzzword for the bin.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;My obsession with predictions in general is a result of trying to identify the right sources to listen to. I am sure it&apos;s not just me constantly on the lookout for trustworthy and reliable content that helps me put my perception of the future into context. And I am also fairly confident that a lot of you, dear readers, are thinking about finding the right direction not just for your businesses, but for yourselves professionally and figure out where your career path is supposed to take you.&lt;/p&gt;
&lt;h4 &gt;The Only Top 10 Game-Changing Sustainable Tech Future Innovation Agile Disruption Strategy Trend Report Workbook Executive Summary CEO Guide 2021 You Need - Download PDF here&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Gartner is one of the first firms releasing their &lt;a href=&quot;https://www.gartner.com/smarterwithgartner/gartner-top-strategic-technology-trends-for-2021/&quot;&gt;Top Strategic Technology Trends for 2021&lt;/a&gt;. The list of topics they&apos;ve considered meaningful enough to make it to the list doesn&apos;t include blockchain! Wow, we&apos;re off to a great start. Sometimes I try to imagine how the season just before Christmas must be really difficult for blockchain enthusiasts and crypto fanboys. Every year they have to sit down and come up with new excuses why the past year was not the breakthrough year for their beloved technology and how the next one is going to be all different with mass-adoption, with unicorn valuations and maybe even with a valid usecase... but I digress.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/career_plan.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Back to Gartner: &quot;This year’s trends fall under three themes: People centricity, location independence and resilient delivery.&quot; - they say. They point out the following nine trends:&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Internet of Behaviors&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Total experience&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Privacy-enhancing computation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Distributed cloud&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Anywhere operations&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Cybersecurity mesh&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Intelligent composable business&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;AI engineering&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Hyperautomation&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;&quot; &gt;For a CEO or COO who would like to enable their workforce to act upon these proclaimed tendencies, this would translate into: process digitalisation, improving access to data, data management processes and preparing for continued organisational readjustments. This is a pretty clear confirmation of the notion that COVID is merely accelerating existing momentums like remote working, data-led decision making and automation in business. What struck me when reading through the report was that data is very much in the center of all emerging trends coming our way, even more dominantly than at first glance. Duh.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But there is an observable shift toward data infrastructure-related investments that enable positive outcomes for the business as a whole instead of just doing machine learning l’art pour l’art to fuel a lighthouse project for AI or to &quot;build a data science muscle&quot;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This means that the nature of innovation is changing by moving away from the shiny and sexy, therefore the allocation of capital has to shift in parallel as a result. So here is a very real question that a lot of engineering managers and data team leaders ask themselves: what kind of competences does my team need in order to be able to deliver on the strategic expectations laid out the business plans for 2021?&lt;/p&gt;
&lt;h4 &gt;Skilling up for 2021&lt;/h4&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;Knowing how to invest today means knowing what talent is needed.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Source: &lt;a href=&quot;https://web.archive.org/web/20201127232546/https://www.elementai.com/news/2020/getting-the-needed-talent-along-the-ai-value-chain&quot;&gt;JF Gagne (CEO and founder of Element AI)&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;What I am seeing is one more pattern that corroborates the growing demand for data engineering talent, and confirms that there is no clear end in sight for it to drop. If the trends listed above turn into tangible initiatives within your company, and you have to start laying out a staffing plan for 2021, there is a very good chance you&apos;ll find the need for someone who understands tech infrastructure, the special nature of data and can deal with the increased amount of alignment meetings.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Maslow&apos;s argument remains true when it comes to my own biased POV,&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;... it is tempting, if the only tool you have is a hammer, to treat everything as if it were a nail.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;But as I always say, don&apos;t believe me, believe the numbers: below just two extracts of the latest researches showing the enormous growth of demand for data engineering skills:&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_occupation_growth.png&quot; alt=&quot;Chart of the fastest growing tech occupations by year-over-year growth, with data engineer first at 50 percent&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Source: &lt;a href=&quot;http://marketing.dice.com/pdf/2020/Dice_2020_Tech_Job_Report.pdf&quot;&gt;Dice Tech Job Report Q1 2020&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_skills.jpg&quot; alt=&quot;Bar chart of the skills most often named in data engineering job ads, led by data engineering, databases and Python&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Source: &lt;a href=&quot;https://opendatascience.com/rise-of-the-data-engineer/&quot;&gt;opendatascience.com&lt;/a&gt;&lt;/p&gt;&lt;/div&gt; 
          &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;I expect more trend reports to come out in the next weeks underlining the need for data professionals in the context of the new normal (WFH, experiences and events going digital etc.) either directly or indirectly. It&apos;s a great sign for everyone who has already embarked on the journey of transitioning their careers into the world of data, and one more reason for rookies to focus more on engineering skills.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The data suggests that data engineering is here to stay, and having observed how data science has evolved as a profession, and especially how first-movers gained an almost unfair advantage compared to whoever tries to get into the game today... it is safe to say that the right time to start data engineering is yesterday. That&apos;s my humble prediction.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #2: Tamás Németh</title>
   <link href="https://dataengineering.academy/idataengineer/2020/11/25/idataengineer-confessions-interview-002.html"/>
   <updated>2020-11-25T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/11/25/idataengineer-confessions-interview-002.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/hMA8EB9jRYg?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Confessions #1: Ramzi Al-Aruri</title>
   <link href="https://dataengineering.academy/idataengineer/2020/11/18/idataengineer-confessions-interview-001.html"/>
   <updated>2020-11-18T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/idataengineer/2020/11/18/idataengineer-confessions-interview-001.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;In this interview series we ask data engineers how they ended up choosing this career and how they see the future for it.&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;You can subscribe to the show on &lt;a href=&quot;https://www.youtube.com/channel/UCzXd80RNUOkANQfcWYw1slg&quot;&gt;Youtube&lt;/a&gt;, &lt;a href=&quot;https://open.spotify.com/show/1hkQwsooJDT7jAq22aMjIh&quot;&gt;Spotify&lt;/a&gt; and &lt;a href=&quot;https://podcasts.apple.com/podcast/idataengineer/id1545671434&quot;&gt;Apple Podcasts&lt;/a&gt; to get notified about new episodes.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&quot;intrinsic&quot; style=&quot;max-width:100%&quot;&gt;&lt;div class=&quot;embed-block-wrapper &quot; style=&quot;padding-bottom:56.20609%;&quot;&gt;&lt;div class=&quot;sqs-video-wrapper&quot; data-provider-name=&quot;YouTube&quot; data-html=&apos;&amp;lt;iframe src=&quot;//www.youtube.com/embed/y99WRhmqyuY?wmode=opaque&amp;amp;enablejsapi=1&quot; height=&quot;480&quot; width=&quot;854&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&amp;gt;&amp;lt;br/&amp;gt;&amp;lt;/iframe&amp;gt;&apos;&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - October 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/11/06/the-data-janitor-letters-october-2020.html"/>
   <updated>2020-11-06T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/11/06/the-data-janitor-letters-october-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://developers.soundcloud.com/blog/testing-sql-for-bigquery&quot;&gt;Testing SQL for BigQuery&lt;/a&gt;&lt;br&gt;&lt;em&gt;Barbara Scherlein, Scala Backend and Data Engineer, Soundcloud&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;When I finally deleted the old Spark code, it was a net delete of almost 1,700 lines of code; the resulting two SQL queries have, respectively, 155 and 81 lines of SQL code; and the new tests have about 1,231 lines of Python code.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.techrepublic.com/article/percona-wants-your-database-to-be-open-source-not-everyone-is-happy-about-it/&quot; target=&quot;_blank&quot;&gt;Why Percona wants your database to be open source, and not everyone is happy about it&lt;/a&gt;&lt;br&gt;&lt;em&gt;Matt Asay, Head of Open Source Strategy and Marketing, AWS&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Percona runs open source databases as managed services, which makes the company popular with customers but less so with competitors.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/swlh/etl-batch-processing-with-kafka-7f66f843e20d&quot; target=&quot;_blank&quot;&gt;ETL Batch Processing With Kafka?&lt;/a&gt;&lt;em&gt;&lt;br&gt;Tomasz Kaszuba, Java Big Data Engineer, Swiss Re&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;For small batch loads using traditional ETL tools is less complicated and much simpler to implement. But if the ETL pipeline needs to handle large amounts of data and scale Kafka wins hands down.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/df-foundation/meet-whale-the-stupidly-simple-data-discovery-tool-9f847c004b47&quot; target=&quot;_blank&quot;&gt;Meet whale! 🐳 The stupidly simple data discovery tool.&lt;/a&gt;&lt;br&gt;&lt;em&gt;Robert Yi, Co-founder &amp;amp; Chief Data Officer, Dataframe&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A Python library that scrapes metadata and formats it as markdown. A Rust CLI interface to search over that data.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.holistics.io/blog/quel-vs-sql/&quot; target=&quot;_blank&quot;&gt;A Short Story About SQL’s Biggest Rival&lt;/a&gt;&lt;br&gt;&lt;em&gt;Cedric Chin, Content Marketing, Holistics Software&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We might have once lived in a world where QUEL and SQL would have continued to duke it out, and where the ‘best’ language might have found its own niches.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20201113182017/https://www.endpoint.com/blog/2020/10/02/postgresql-binary-search-correlated-data-cte&quot; target=&quot;_blank&quot;&gt;Using CTEs to do a binary search of large tables with non-indexed correlated data&lt;/a&gt;&lt;br&gt;&lt;em&gt;David Christensen, Senior Software and Database Engineer, End Point&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The initial query went from timing out in the webservice in question to returning results in a fraction of a second with basic binary search.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.charlesfarina.com/new-bigquery-integration-for-app-web-properties/&quot; target=&quot;_blank&quot;&gt;New BigQuery Integration for GA4 Properties&lt;/a&gt;&lt;br&gt;&lt;em&gt;Charles Farina, Head of Innovation, Adswerve, Inc.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Free BigQuery export from GA to all customers.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://insights.project-a.com/how-to-set-up-a-multi-touch-attribution-model-a24c4d8e6b8d&quot; target=&quot;_blank&quot;&gt;How to set up a multi-touch attribution model&lt;/a&gt;&lt;br&gt;&lt;em&gt;Cyprien Marcos, Business Intelligence Manager, Project A Ventures&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We show you how you can easily set up a multi-touch attribution model to track website conversions with Google Analytics, Google Tag Manager and a Jupyter notebook.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>How to apply to the bootcamp</title>
   <link href="https://dataengineering.academy/2020/10/17/data-engineer-how-to-apply.html"/>
   <updated>2020-10-17T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/10/17/data-engineer-how-to-apply.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;I firmly believe that within the legal and moral framework of our society, education is by far the most potent transformational force that enables socioeconomic mobility. Therefore, we&apos;re working hard to make our bootcamp easily accessible for everyone, and to help our participants unlock professional and economic opportunities though education. In other words, we teach you something that has the potential to take you to the next level in your professional career.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;In an ideal world, &lt;a href=&quot;/&quot;&gt;Pipeline Academy&lt;/a&gt; would have a suitable offering for everyone interested in data engineering regardless of geography, existing skills, preferred learning method etc., but we&apos;re not there yet. But how do you measure access to a school, and how do you communicate that?&lt;/p&gt;
&lt;h4 &gt;Un/Restricted access&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If you take a look at the key performance indicators that educational institutions like to use to showcase their superiority, you might stumble upon some organisations priding themselves in how low their acceptance rates are (e.g. &quot;&lt;a href=&quot;https://www.prepscholar.com/sat/s/colleges/Harvard-admission-requirements&quot;&gt;only 4.7% of applicants are accepted to Harvard&lt;/a&gt;&quot;). When high demand faces scarcity, without intervention the result will inevitably be an increased price. Approaching it from an economic point of view, it only makes sense for a business to capitalise on the reputation it built for itself through years of impactful outcomes (successful graduates), and turn its offering into a product of exclusivity. But let&apos;s just take a step back: &lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Is our society across the board benefitting from making the highest quality education a luxury good?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Is building fences around knowledge not counterproductive to our shared welfare and reducing our future opportunities for the sake of short-term profits?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Is this trend not solidifying socioeconomic inequality?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;My point: you&apos;re never going to see us pride ourselves in excluding people. A low acceptance rate at a bootcamp should only be acceptable if there is no way of scaling the number of seats without sacrificing the quality of the service. Otherwise it is a likely indicator of people trying to cash in, instead of trying to help others. Preselection is absolutely necessary, but it seems like we&apos;ve forgotten why we&apos;re doing it, and today it only serves schools instead of students.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_bootcamp_how_to_apply.jpg&quot; alt=&quot;A queue of people waiting outside a shop&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;How to apply&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We&apos;ve designed our &lt;a href=&quot;/apply/&quot;&gt;application process&lt;/a&gt; with one principle in mind: how can we make sure that we include everyone who is seeking opportunity and willing to put in the work. The steps don&apos;t include anything unexpected, it&apos;s more about how we deal with our applicants:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;We&apos;re doing our best to stay easily approachable so the folks who are interested in joining Pipeline Academy can talk to us, ask questions and get a good feeling for what to expect. You can &lt;a href=&quot;https://calendly.com/pipeline-peter&quot;&gt;schedule a call right here&lt;/a&gt; anytime even without submitting an application.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Communicating &lt;a href=&quot;/faq/&quot;&gt;the prerequisites of our bootcamp&lt;/a&gt; is supposed to help you figure out what you need to do in order to get admitted, not to discourage you. We would like to be very clear on what is expected from you before and during the program, and make sure you get some initial guidance before you get started.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;After filling out the &lt;a href=&quot;/apply/&quot;&gt;application form&lt;/a&gt; on our website you&apos;ll receive an email as a response (potentially with some answers to your questions you&apos;ve shared via the form) that includes a link so you can select a slot in our calendar for a call that usually takes about 20-30 minutes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;During the admissions call we like to learn about your motivation and attitude besides your professional experience: our goal is to understand why and how you can make this happen instead of comparing you to other candidates. Usually we ask you about the organisations and the teams you&apos;ve worked with in the past, the technologies you are familiar with and some projects you were involved in. We don&apos;t do any coding tests (yet) though. We get back to you within a couple of days after the call with the decision about your admission.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;We will always encourage everyone to learn more, even if this means that you&apos;re not going to be spending money with us. We are happy to point you towards online programs or other alternatives to our bootcamp if we can&apos;t admit you for some reason. Again, the goal is to help you on your path of becoming a data engineer. Period.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Our school and your peers remain accessible for you even after the bootcamp. You&apos;ll get invited to our events if you want to, and we help you with your job search even well after the training.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_bootcamp_admissions.jpg&quot; alt=&quot;A participant raising a hand in class&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;That&apos;s pretty much it. There is no magic, no secret test, no tricky coding challenge to filter you out. Nevertheless, the expectations are high, and we promise that your time at Pipeline Academy won&apos;t be boring.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Other actions we proudly take to ensure everybody has a chance to get closer to the craft of data engineering and our coding bootcamp:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Communicating our approach and actions in a very transparent manner,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Design our curriculum to be accessible for applicants with different levels of experience and skills,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Guiding rookies just as experienced professionals who approach us with questions,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Supporting the data engineering community actively with events and professional input in various formats,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Offering multiple financing options and in some cases even scholarships,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;…&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;If you consider applying to Pipeline Academy, we encourage you to do your research (&lt;a href=&quot;/blog.html&quot;&gt;the blog&lt;/a&gt; is a great starting point) and check if data engineering is a career path you&apos;d like to embark upon. If you feel motivated but have some doubts, just talk to us to see if there is a way we can support you on your journey.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - September 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/10/13/the-data-janitor-letters-september-2020.html"/>
   <updated>2020-10-13T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/10/13/the-data-janitor-letters-september-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20201021082322/https://counting.substack.com/p/data-cleaning-is-analysis-not-grunt&quot; target=&quot;_blank&quot;&gt;Data Cleaning IS Analysis, Not Grunt Work&lt;/a&gt;&lt;br&gt;&lt;em&gt;Randy Au, Quantitative UX Researcher, Google&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Also, most data cleaning articles suck.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.koehntopp.info/2020/09/26/my-private-data-warehouse.html&quot; target=&quot;_blank&quot;&gt;Importing account statements and building a data warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Kristian Köhntopp, Senior Scalability Engineer, &lt;/em&gt;&lt;a href=&quot;http://booking.com/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Booking.com&lt;/em&gt;&lt;/a&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I was experimenting with importing the account statements from my German Sparkasse, which at that time were being made available as a CSV.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.ebayinc.com/engineering/ou-online-analytical-processing/&quot; target=&quot;_blank&quot;&gt;Our Online Analytical Processing Journey with ClickHouse on Kubernetes&lt;/a&gt;&lt;br&gt;&lt;em&gt;Sudeep Kumar, Member of Technical Staff - 2, ebay&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;With our new, cross-region aware OLAP pipeline, we reduced our overall infrastructure footprint by over 90 percent.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/swlh/dawn-of-dataops-can-we-build-a-100-serverless-etl-following-ci-cd-principles-3ca587ba1ec0&quot; target=&quot;_blank&quot;&gt;Dawn of DataOps: Can We Build a 100% Serverless ETL Following CI/CD Principles?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Luis Velasco, Cloud Data Analytics Specialist, Google&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Is it time to enjoy the benefits of DevOps in the informational space?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20201010062257/https://kislayverma.com/programming/publish-events-not-logs/&quot; target=&quot;_blank&quot;&gt;Publish events, not logs&lt;/a&gt;&lt;br&gt;&lt;em&gt;Kislay Verma, Software Engineer, Cure.Fit&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I believe that logging as understood commonly is an ad hoc activity, and cannot handle the unknown-unknowns of a production system. We need to switch to an event perspective to leverage it more effectively for system design and reliability.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/expedia-group-tech/be-vigilant-about-time-order-in-event-based-data-processing-cbfde600dd7d&quot; target=&quot;_blank&quot;&gt;Be Vigilant about Time Order in Event-Based Data Processing&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mingwei Li, Senior Software Engineer, Expedia Group Technology&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;How to handle the timing of events.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineer Job Description Boilerplates</title>
   <link href="https://dataengineering.academy/2020/10/12/data-engineer-job-description-boilerplates.html"/>
   <updated>2020-10-12T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/10/12/data-engineer-job-description-boilerplates.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Are you an HR-professional recruiting new members for your organisation’s data team? If you are looking for talent to fill a data engineer role, you have probably faced the challenge of writing a proper job description. Your CTO/Director of Engineering/Head of Data is very busy, and gave you only some high-level points, but nothing that fits into a JD. Don&apos;t worry, you are not alone.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Below you can find some boilerplate wording for the &apos;Tasks &amp;amp; Responsibilities&apos; section of a data engineer job description. Adjust, adapt and complete it the way you see fit, the point is really to make you think less about the wording and more about finding the right person for the job. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;One hint: be selective in what you write into the job ad, more does not equal better in this case.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_job_description.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Data engineer job description sample text&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Work closely with the product team on the platform&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Develop and scale the recommendation engine and comparable algorithms&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Interface with the infrastructure team&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Scale the architecture to manage increased traffic&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Work on the backend structure, the API and partner integrations&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Start integrating the recommendation engine&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Measure and ensure data quality across the organisation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Use agile software development processes to deliver features&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Build  and document solid data pipelines&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Clean, transform, and aggregate data from different sources&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Develop and document processes for data mining, data modeling and data warehousing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Support building complex algorithms that provide business value to the customers&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Ingest and aggregate data from both internal and external data sources&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Help streamline the data science processes&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Plan data models and architecture&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Implement customer lifecycle and retention models based on an existing methodology&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Employ an array of coding languages and tools to set up the data infrastructure&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Interface with the reporting team and support their objectives with infrastructure tweaks&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Execute/implement new data product features end-to-end&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Engage with the machine learning team and uncover hidden efficiencies in the pipelines&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Turn high-volume activity data into a highly accessible resource&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Ingest and aggregate data from both internal and external data sources&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Transform the available data into meaningful insights working with the analyst team&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Develop simple models and integrate them with the visualisation tools&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Work closely with the data science and business intelligence teams to develop data models and pipelines for research, reporting, and machine learning&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Integrate state-of-the-art data management and software engineering technologies&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Tap into new data streams from third-party APIs&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Create custom software components for the data platform&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Collaborate with the stakeholders in an interdisciplinary team&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Research new approaches for making data accessible to customers&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Connect legacy and new data systems together&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Define basic tooling, metrics and other solutions and maintain quality to the highest standards&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Write, extend and debug microservices for planned features&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Collaborate with other engineers, data analysts, and product managers&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Lead data strategy and inform the product strategy team on how world-class data experiences are built&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Take leadership opportunities and shape the data culture within the organisation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Build near real-time and batch data processing pipelines&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Design and optimise low latency systems&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Build highly reliable but flexible service infrastructure&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Maintain tooling and enable algorithms to move into production faster&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Relate and match entities from different data streams&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Increase the implementation speed of data tools&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Creating secure processes to keep the data pipelines safe&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Improve data models and foster data-driven decision making&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Draw a comprehensive picture of user flows and enable deeper analysis&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Specify data requirements and pre-processing routines&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Hold companywide data trainings&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Develop solutions for automatic labeling of data&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Model front end and back end data sources&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Design and optimise complex queries and deliver user value&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Work closely with data scientists on modeling&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Grow the data competency across the company &lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Presenting and Teaching via Zoom 101</title>
   <link href="https://dataengineering.academy/2020/10/03/presenting-and-teaching-via-zoom-101.html"/>
   <updated>2020-10-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/10/03/presenting-and-teaching-via-zoom-101.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Yes, we all miss the full bandwidth of reality, but let&apos;s face it, we&apos;re left with experiencing the world via videoconferencing tools for now.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;br&gt;You can find dozens of videos answering specific questions related to the setup, the tools and the nitty-gritty, here I summarize what I figured -- standing on the shoulders of giants, as always. (You don&apos;t want to know how many different items and setups I tried before settling on the bare minimum.)&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;Rule #1: no wireless.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Wireless technology is on the brinks on dealing with this device density given the bandwidth available. If you want a reliable setup, use cables. Most importantly use an Ethernet cable to connect your computer to the router. This is the 80%.&lt;/p&gt;
&lt;h4 &gt;Presenting&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This is the more focused, defined format so figure the following:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;camera distance: you should fit in the screen, almost fill the frame, but being able to use all your gestures,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;lens height: get the lens at eye level, your nostrils are not that fun,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;light: have a window near with natural light, preferably on the side, for extra fun you can add a 5600K bulb in a lamp, that bounces its light on the wall or coming at your face at an angle,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;background: should be something neutral, a bookshelf is great, and Zoom&apos;s virtual background with a green screen actually runs on a 5 year old computer,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;sound: generic Apple Earpods will just do fine, but getting an entry level lavalier mic would not hurt,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;camera: the built-in Mac camera is okay. If you talk to a monitor, don&apos;t go further than the &lt;a href=&quot;https://www.idealo.de/preisvergleich/OffersOfProduct/3070374_-hd-pro-c920-logitech.html&quot;&gt;Logitech C920&lt;/a&gt;,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;a presenter — hard to get a wired one.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;This is  the backstage. In this specific setup I used an old &lt;a href=&quot;https://web.archive.org/web/20200929080123/https://joby.com/us-en/gorillapod-flexible-camera-tripods/&quot;&gt;Joby Gorillapod&lt;/a&gt; to fix the &lt;a href=&quot;https://amazon.de/-/en/gp/product/B07BT2JYH4/&quot;&gt;Neewer LED lamp&lt;/a&gt; with 5600K.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/2020-09-22-14.53.22.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;This is what I looked like presenting online at &lt;a href=&quot;https://budapestdata.hu/2020/en/&quot;&gt;Budapest Data Forum 2020&lt;/a&gt;. &lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/screenshot_2020-09-23_at_10.26.24.png&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Teaching&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Different styles and methods ask for different setups, still I believe this a minimum viable approach:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;real estate: more of it! You will need an extra big monitor to be able to follow the interaction, the chat, the faces, your slides and your terminal,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;drawing: an 6th generation iPad with a first generation Apple Pencil as a separate Zoom user will do the trick if you want to draw on your slides or just simply use it as a whiteboard,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;non-technical items: I let you figure the ones I found useful for teaching — just by looking at the photo :)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/2020-05-18-08.46.47.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;This is what I looked like live teaching via Zoom during the &lt;a href=&quot;/2020/07/08/this-is-not-a-test-this-is-a-summer-camp-part-2.html&quot;&gt;Summer Camp&lt;/a&gt;. &lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/daniel_the_data_janitor-5l5r5n.png&quot; alt=&quot;Daniel Molnar, the Data Janitor&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;Postscript&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Choosing the software we have to use is not always on us. Quality-wise Facetime, Bluejeans and Zoom are the descending order in our experience. Depending on your computing power you might want to give a chance to &lt;a href=&quot;https://web.archive.org/web/20200929003848/https://snapcamera.snapchat.com/&quot;&gt;Snap Camera&lt;/a&gt;, especially &lt;a href=&quot;https://www.snapchat.com/unlock/?metadata=01&amp;amp;type=SNAPCODE&amp;amp;uuid=16839bd69c67492696d6ccf1296ad31e&quot;&gt;Meeting Gestures by Cameron Hunter&lt;/a&gt;, or if you have a power station give a chance to &lt;a href=&quot;https://www.mmhmm.app/&quot;&gt;mmhmm&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Data Engineering Podcast With The Data Janitor</title>
   <link href="https://dataengineering.academy/2020/09/23/data-engineering-podcast-pipeline-academy.html"/>
   <updated>2020-09-23T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/09/23/data-engineering-podcast-pipeline-academy.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Tobias Macey, the Joe Rogan of the data engineering scene interviewed Daniel in the Data Engineering Podcast. &lt;a href=&quot;https://www.dataengineeringpodcast.com/pipeline-data-engineering-academy-episode-151/&quot;&gt;Give it a listen right here!&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.dataengineeringpodcast.com/pipeline-data-engineering-academy-episode-151/&quot;&gt;Tobias’  summary&lt;/a&gt; of the insightful conversation in episode #151:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Data engineering is a constantly growing and evolving discipline. There are always new tools, systems, and design patterns to learn, which leads to a great deal of confusion for newcomers. Daniel Molnar has dedicated his time to helping data professionals get back to basics through presentations at conferences and meetups, and with his most recent endeavor of building the Pipeline Data Engineering Academy. In this episode he shares advice on how to cut through the noise, which principles are foundational to building a successful career as a data engineer, and his approach to educating the next generation of data practitioners. This was a useful conversation for anyone working with data who has found themselves spending too much time chasing the latest trends and wishes to develop a more focused approach to their work.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;About the host and the show (taken from &lt;a href=&quot;https://web.archive.org/web/20200930155323/https://www.dataengineeringpodcast.com/about/&quot;&gt;the website of the podcast&lt;/a&gt;):&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;“Tobias Macey is a dedicated engineer with experience spanning many years and even more domains. He currently manages and leads the Technical Operations team at MIT Open Learning where he designs and builds cloud infrastructure to power online access to education for the global MIT community. He also owns and operates &lt;/em&gt;&lt;a href=&quot;https://www.boundlessnotions.com/&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Boundless Notions, LLC&lt;/em&gt;&lt;/a&gt;&lt;em&gt; where he offers design, review, and implementation advice on data infrastructure and cloud automation.”&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;“The Data Engineering Podcast tackles a new approach to data management every week. Each new episode provides useful and informative insights into the projects, platforms, and practices that data engineers, team leaders, and data scientists need to know about to learn and grow in their career.”&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Like and subscribe on &lt;a href=&quot;https://podcasts.apple.com/us/podcast/data-engineering-podcast/id1193040557?mt=2&quot;&gt;Apple Podcasts&lt;/a&gt; and give the show a follow on &lt;a href=&quot;https://twitter.com/DataEngPodcast&quot;&gt;twitter&lt;/a&gt;!&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_podcast.jpg&quot; alt=&quot;Data Engineering Podcast logo&quot;&gt;
</content>
 </entry>
 
 <entry>
   <title>Data engineer salaries in Germany 2020</title>
   <link href="https://dataengineering.academy/2020/09/22/data-engineer-salary-germany-2020.html"/>
   <updated>2020-09-22T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/09/22/data-engineer-salary-germany-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;TL;DR: the expected annual gross salary for a junior or mid-level data engineer in Germany is between €45.000-75.000 with a steep increase after a couple of years in the field. &lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Below a rundown done in August 2020 based on secondary sources from across the internet. This has been verified by our primary research with stakeholders of the Berlin data ecosystem. (Note: if you don’t like the inconsistent structure of the below spreadsheet, you are not alone.)&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;Data Engineer Salary Germany 2020&lt;/h4&gt;
&lt;/div&gt;
&lt;style type=&quot;text/css&quot;&gt;
  .tg  {border-collapse:collapse;border-spacing:0;}
  .tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
    overflow:hidden;padding:10px 5px;word-break:normal;}
  .tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
    font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}
  &lt;/style&gt;
&lt;table class=&quot;tg&quot;&gt;
&lt;thead&gt;
  &lt;tr&gt;
    &lt;th class=&quot;tg-58we&quot;&gt;Source&lt;/th&gt;
    &lt;th class=&quot;tg-sdg2&quot;&gt;Position&lt;/th&gt;
    &lt;th class=&quot;tg-sdg2&quot;&gt;Salary&lt;/th&gt;
  &lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-a3en&quot; rowspan=&quot;3&quot;&gt;&lt;a href=&quot;https://datadrivencompany.de/data-engineer-beschreibung-aufgaben-tools-gehalt/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;datadrivencompany.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Junior Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;€40.000-60.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Mid-level Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;50.000-90.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Senior Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;           €&lt;/span&gt;80.000-120.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-a3en&quot; rowspan=&quot;3&quot;&gt;&lt;a href=&quot;https://orange-quarter.com/data-analytics-salaries-berlin-2020/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;orange-quarter.com&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Junior Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;40.000-50.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Mid-level Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;50.000-65.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Senior Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;65.000-90.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-a3en&quot; rowspan=&quot;2&quot;&gt;&lt;a href=&quot;https://datasolut.com/data-engineer-berufsbild/&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;datasolut.com&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Junior Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;50.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;Mid-level Data Engineer&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;70.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-vpq9&quot;&gt;&lt;a href=&quot;https://www.hays.de/jobprofile/data-engineer&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;hays.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;Data Engineer (level of seniority unspecified)&lt;/span&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;45.000-70.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-vpq9&quot;&gt;&lt;a href=&quot;https://de.talent.com/jobs?k=data+engineer&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;neuvoo.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;Data Engineer (level of seniority unspecified)&lt;/span&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;70.750&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-vpq9&quot;&gt;&lt;a href=&quot;https://www.glassdoor.de/Geh%C3%A4lter/data-engineer-gehalt-SRCH_KO0,13.htm&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;glassdoor.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;Data Engineer (level of seniority unspecified)&lt;/span&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;60.000&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-vpq9&quot;&gt;&lt;a href=&quot;https://www.stepstone.de/gehalt/Data-Engineer.html&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;stepstone.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;Data Engineer (level of seniority unspecified)&lt;/span&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;54.600&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
    &lt;td class=&quot;tg-vpq9&quot;&gt;&lt;a href=&quot;https://www.gehalt.de/beruf/data-engineer&quot; target=&quot;_blank&quot; rel=&quot;noopener noreferrer&quot;&gt;gehalt.de&lt;/a&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-ejvj&quot;&gt;&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;Data Engineer (level of seniority unspecified)&lt;/span&gt;&lt;/td&gt;
    &lt;td class=&quot;tg-hf04&quot;&gt;
&lt;span style=&quot;font-weight:400;font-style:normal&quot;&gt;€&lt;/span&gt;62.777&lt;/td&gt;
  &lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;&lt;em&gt;Click the links in the Source column to see the the referenced website. Please note, the above information is subject to change.&lt;/em&gt;&lt;br&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Things you should know when looking at these figures:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Expected salaries are not advertised in job listings in Germany, and most companies are highly secretive about this question which adds a lot of inefficiencies to the job search.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The listed salaries are before taxes, which are determined by various factors. Here&apos;s &lt;a href=&quot;https://www.brutto-netto-rechner.info/&quot;&gt;one of the websites&lt;/a&gt; helping you calculate what your net salary expectations should look like.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;There are considerable regional differences in cost of living within Germany, make sure to adjust the above figures accordingly. Also, you&apos;ll find that there is a major mismatch between what a data engineer makes in Silicon Valley compared to Europe in general: the job markets, tax structures and the cost of living are just so different, that a comparison hardly makes sense.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;There is an upwards trend in salary average and the number of open positions for engineering roles. It should also be noted, that salaries increase with a higher level of seniority much faster than for most other positions, and that the starting salaries for more commoditized engineering positions (e.g. frontend developer) are usually significantly lower than for data engineering roles.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineer_salary.jpg&quot; alt=&quot;The Brandenburg Gate seen along an empty Berlin street&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt; Some considerations for your job search:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Experience with some specific tools or programming language, background in a certain industry, prior experience in coding or even with managing people have a strong positive impact on salary expectations. Use that when negotiating.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The scope of the data engineer&apos;s role is broad, which is one of the reasons why the position often runs under a different name: machine learning engineer, BI engineer, backend engineer data platforms, data lead etc., furthermore some companies are looking for engineers for using specific tools: Spark engineer, Hadoop engineer etc.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The career path is also not straightforward in terms of naming conventions for more senior positions, so senior data engineer, data architect, big data architect, head of data etc. are also keywords one should check.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Do your research: check out local job platforms, job metasearch engines, &lt;a href=&quot;https://www.glassdoor.de/&quot;&gt;glassdoor&lt;/a&gt; and &lt;a href=&quot;https://www.kununu.com/&quot;&gt;kununu&lt;/a&gt; for company reviews (these platforms can save you from some employer-induced PTSD), and make sure to also read the latest news about the organisation you are considering joining.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Salary is not everything: when &lt;a href=&quot;https://365datascience.com/data-engineer-interview-questions/&quot;&gt;interviewing&lt;/a&gt;, ask about team structure and who you are supposed to report to, career paths, stack and learning opportunities. Having a strong mentor can bring you more value on the long run than a few more euros.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Keep on learning, but even more importantly: keep on working. Create a hands-on portfolio that stands out, experiment with tools and contribute to open source projects to show your skills.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Data engineering is an accessible career path that involves competencies that most career advisors and recruiters would consider essential for the upcoming decades. Regardless if you are a generalist or a specialist, working in tech or any other sector, understanding the flow of data and how to manage it makes any employee into an asset. &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - August 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/09/16/the-data-janitor-letters-august-2020.html"/>
   <updated>2020-09-16T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/09/16/the-data-janitor-letters-august-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20201020121938/https://dataform.co/blog/the-startup-data-stack-starter-pack&quot;&gt;The startup data stack starter pack (2020)&lt;/a&gt;&lt;br&gt;&lt;em&gt;Lewis Hemens, CTO, Dataform&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I advise a lot of people on how to build out their data stack, from tiny startups to enterprise companies that are moving to the cloud or from legacy solutions. There are many choices out there, and navigating them all can be tricky.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://highgrowthengineering.substack.com/p/why-is-dbt-so-important-&quot; target=&quot;_blank&quot;&gt;Why is dbt so important?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Stephen Whitworth, Senior Software Engineer, Monzo Bank&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If you choose not to use dbt, you’ll probably waste time building a less-fully featured, buggy implementation of it yourself. Give it a serious look.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@rahulj51/guiding-principles-for-a-data-engineering-team-7baaa35b2954&quot; target=&quot;_blank&quot;&gt;Guiding principles for a data engineering team&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/rahulj51&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Rahul Jain&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Principal Engineering Manager, Data engineering and BI platform, Omio/GoEuro&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We prefer boring but battle tested technologies over tech-fetishism.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://altinity.com/blog/clickhouse-and-redshift-face-off-again-in-nyc-taxi-rides-benchmark&quot; target=&quot;_blank&quot;&gt;ClickHouse &amp;amp; Redshift Face Off in NYC Taxi Rides Benchmark&lt;/a&gt;&lt;br&gt;&lt;em&gt;Alexander Zaitsev, Co-founder, Altinity&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;2020 versions of both ClickHouse and Redshift show much better performance. However, open source ClickHouse continues to outperform Redshift on similarly sized hardware, and the difference increases as the query complexity grows.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://eugeneyan.com/writing/end-to-end-data-science/&quot; target=&quot;_blank&quot;&gt;Unpopular Opinion - Data Scientists Should Be More End-to-End&lt;/a&gt;&lt;br&gt;&lt;em&gt;Eugene Yan, Applied Scientist, Amazon&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Going out of the regular DS &amp;amp; ML job scope helped with delivering more value, faster.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@steve.yegge/dear-google-cloud-your-deprecation-policy-is-killing-you-ee7525dc05dc&quot; target=&quot;_blank&quot;&gt;Dear Google Cloud: Your Deprecation Policy is Killing You&lt;/a&gt;&lt;br&gt;&lt;em&gt;Steve Yegge, Head Dude, Ghost Track&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Backwards compatibility keeps systems alive and relevant for decades.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.shreya-shankar.com/ai-saviorism/&quot; target=&quot;_blank&quot;&gt;Get rid of AI Saviorism&lt;/a&gt;&lt;br&gt;&lt;em&gt;Shreya Shankar, Machine Learning Engineer, Viaduct&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Machine learning is a tool, not a panacea. If our tools don’t immediately work for them, it’s not their fault.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The state of coding bootcamps in 2020</title>
   <link href="https://dataengineering.academy/2020/09/03/the-state-of-coding-bootcamps-in-2020.html"/>
   <updated>2020-09-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/09/03/the-state-of-coding-bootcamps-in-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;My role at Pipeline  Academy entails a variety of things I really love to do, one of them is setting up a strategy for the company and translating it into actionable operational steps.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s pretty much like using any map:&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Understanding the industry/market that my company is active in by analysing where we came from and where we are right now.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Define a desirable future state or an idea of the value you generate (use your preferred framework).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;After that comes the fun part, finding the path that gets us there.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;&quot; &gt;So here is my step #1, let&apos;s see how the world of coding bootcamps looks like from different perspectives (industry analyst, bootcamp organiser, participant) based on some of the key resources I&apos;ve encountered during my never-ending voyage on the internet. If you&apos;d like to get closer to understanding this fairly new educational format and you are looking for a non-comprehensive starting point, you might find the links below useful.&lt;/p&gt;
&lt;h3 &gt;Field guide &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;One of the most important write-ups about coding bootcamps has been done by Quincy Larson, the teacher who founded &lt;a href=&quot;http://freecodecamp.org&quot;&gt;freeCodeCamp.org&lt;/a&gt;: this very no-nonsense document called &lt;a href=&quot;https://www.freecodecamp.org/news/coding-bootcamp-handbook/&quot;&gt;The Coding Bootcamp Handbook&lt;/a&gt; is a reality check and golden piece of advice for anyone who just recently fell in love with the idea of going to a bootcamp and magically upgrade their lives &apos;get-rich-quick-scheme&apos;-style. The sobering reality is:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Going to a bootcamp can be the best decision you ever made. Or it can be an awkward financial setback. Do your research. Save up your money. Learn coding fundamentals first. Bootcamps aren&apos;t magic. They aren&apos;t going to do the work for you. In the end, the experience is what you make of it. So make the most of it.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;The DOs and DON&apos;Ts of this universe have been shared in a &lt;a href=&quot;https://web.archive.org/web/20200208114700/http://blog.thefirehoseproject.com:80/posts/coding-bootcamp-strategy/&quot;&gt;four-part story&lt;/a&gt; by Marco Morawec, Co-Founder of the Firehose Project that has meanwhile been acquired by Trilogy education (that has meanwhile been acquired by 2U Inc.). It presents solid evidence-based guidance for the makers of bootcamps and for aspiring participants as well, and makes you cut the bullshit and focus on what brings you forward:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Quality and depth of the curriculum  &lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Emphasis on the coding fundamentals – algorithms and data structures&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;A bootcamp team that is technical and 100% focused on student outcomes, not marketing metrics&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Everything else a coding bootcamp focuses on is useless”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Here is the rub: whether you are planning to organise a coding bootcamp or you are considering participating in one, make sure you keep it real. While there are way too many buzzwords and promises floating around, what you put in is what you&apos;ll get out of it.&lt;/p&gt;
&lt;h3 &gt;The coding bootcamp market&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Most of the public information about the state of the market is based on data from the US and Canada. While the formal education system in Europe is different in many aspects, and it seems like there is very little structured and comparative data available on the current size and quality of the industry, the European countries face similar challenges in filling the competence-gap on the highly competitive and fast-paced job market. Therefore, any implications and assumptions need to be treated with caution.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The media loves to reference the latest market analysis by &lt;a href=&quot;http://coursereport.com&quot;&gt;CourseReport.com&lt;/a&gt;, since this coding bootcamp review platform offers its readers the most basic, and at the same time the most crucial key performance indicators in a simple format. In the comprehensive &lt;a href=&quot;https://www.coursereport.com/coding-bootcamp-ultimate-guide&quot;&gt;Coding Bootcamps in 2020: Your Complete Guide To The World of Bootcamps&lt;/a&gt; report you get access to a lot of details about their research and the most important findings, that they&apos;ve summarised like this (note: the data is self-reported by the participating schools):&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Coding bootcamps are intensive, accelerated learning programs that teach beginners digital skills like Full-Stack Web Development, Data Science, Digital Marketing, UX/UI Design, Cybersecurity, and Technical Sales.  &lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;The average bootcamp costs ~$13,500, and graduates report an average starting salary of $67,000. Bootcamps can vary in length from 6 to 28 weeks, although the average bootcamp is ~14 weeks long.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Coding bootcamps teach modern, widely used programming languages and frameworks like Ruby on Rails, Python on Django, JavaScript, and PHP stacks through project-based learning. Students graduate from bootcamps with a portfolio, an online presence, interview skills and more. Most bootcamps help graduates find an internship or match students with an employer network – in fact, in Course Report&apos;s most recent research, 83% of bootcamp alumni report being employed in programming jobs.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;There are bootcamp campuses in over 85 cities throughout the US/Canada. Coding bootcamps are predicted to graduate 23,000 students and gross $309MM in tuition revenue in 2019.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;The average bootcamper has 6 years of work experience, has at least a Bachelor&apos;s degree, and has never worked as a programmer. However, the number of students with degrees appears to be declining slightly over time.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;A lot has changed since bootcamps launched in 2012: University Bootcamps are now competing with household bootcamp names, payment options like Income Share Agreements and Deferred Tuition have exploded in recent years, and many bootcamps are dipping into the corporate training market.”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;&quot; &gt;CourseReport started releasing its &lt;a href=&quot;https://www.coursereport.com/reports/2017-coding-bootcamp-market-size-research&quot;&gt;Coding Bootcamp Market Size Study in 2017&lt;/a&gt;, and since then they have done it in &lt;a href=&quot;https://www.coursereport.com/reports/2018-coding-bootcamp-market-size-research&quot;&gt;2018&lt;/a&gt; and &lt;a href=&quot;https://www.coursereport.com/reports/coding-bootcamp-market-size-research-2019&quot;&gt;2019&lt;/a&gt; as well. In addition, a separate &lt;a href=&quot;https://www.coursereport.com/reports/coding-bootcamp-job-placement-2018&quot;&gt;Coding Bootcamp Alumni Outcomes &amp;amp; Demographics Report has been issued 2018&lt;/a&gt; and &lt;a href=&quot;https://www.coursereport.com/reports/coding-bootcamp-job-placement-report-2019&quot;&gt;2019&lt;/a&gt; to share more transparency for (potential) students and manage expectations towards promised outcomes.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/coding_bootcamp.png&quot; alt=&quot;Infographic on the growth of coding bootcamps: 2019 market size and market growth from 2013 to 2019&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;An even more academic approach can be found in the RTI Press publication &lt;a href=&quot;https://www.rti.org/rti-press-publication/alternative-and-independent/fulltext.pdf&quot;&gt;Alternative and Independent: The Universe of Technology Related &quot;Bootcamps&quot;&lt;/a&gt; by Caren A. Arbeit, Alexander Bentz, Emily Forrest Cataldi and Herschel Sanders. This 2017 report does a great job in highlighting the reasons behind the critique that feels almost like an &lt;a href=&quot;https://evolllution.com/revenue-streams/market_opportunities/the-problem-with-bootcamps-research-uncovers-transparency-issues/&quot;&gt;overused cliche&lt;/a&gt; in 2020:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“While the bootcamp model holds a great deal of promise, it is currently an unregulated, unaccredited sector of higher education with little transparency and no independent outcomes data.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;My favourite and most comprehensive market research has been released on CareerKarma: written by James Gallagher, the &lt;a href=&quot;https://careerkarma.com/blog/bootcamp-market-report-2020/&quot;&gt;State of the Bootcamp Market Report 2020&lt;/a&gt; is a highly professional and extensive analysis. It is an attempt to cover all the aspects that belong in a proper market exploration while trying to redefine some of the standards of measuring the dynamics of this specific market (note: &lt;a href=&quot;https://cirr.org/data&quot;&gt;CIRR&lt;/a&gt; is one attempt to move towards a clear standard in reporting outcomes). He ends up with a significantly higher gross total market revenue for 2019 than CourseReport, which shows how intransparent the US market is - not to mention the rest of the world. In addition, the report includes a map of the ISA landscape, which will be subject to consolidation in the upcoming years similarly to the bootcamp business.&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“If bootcamps continue to innovate and find success equipping people with tech industry-ready skills, the bootcamp sector will likely see significant growth over the next five years — both in the number of bootcamps and the number of students enrolling in them. There also may be a period of consolidation in the space, as existing players become more dominant and look to expand their influence. In addition, the growth of coding bootcamps and the increased diversity in their offerings suggests that the bootcamp model could have wider applications than solely within the context of programming instruction. As courses in digital marketing, technology sales, and other non-coding technical disciplines have become more common, it may be a signal that bootcamps are looking to expand outside of programming disciplines. This would allow the bootcamp model to be applied to a much larger market segment: post-secondary training in non-technical fields. Overall, bootcamps have proven themselves an effective pathway to a career in tech for thousands of people, and all signs indicate that the sector will continue to grow in coming years.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Jay Wengrow, the CEO of a coding bootcamp called Actualize &lt;a href=&quot;https://www.linkedin.com/pulse/state-coding-bootcamps-2018-jay-wengrow/&quot;&gt;summarizes the state of the industry (as of 2018)&lt;/a&gt; based on his insightful experiences building up his company focusing on the challenges of the past years and the ones to come in the near future.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Organic growth is clearly visible, however not at the pace of software startups. Every time you encounter new data about the market or participants, you have to double-check the source and be cautious when basing decisions on them.&lt;/p&gt;
&lt;h3 &gt;Business model and scaling&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;The latest analysis called &lt;a href=&quot;https://www.classcentral.com/report/bootcamps-and-isas/&quot;&gt;Bootcamps and ISAs: Economics, Challenges, and Opportunities&lt;/a&gt; by Christoph Rindlisbacher looks at the coding bootcamp concept and business model from the perspective of a financial investor and comes to a conclusion that won&apos;t please VCs, but it ultimately shows a positive outcome, highlighting that the last decade was fairly turbulent:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“The structural obstacles to scaling a bootcamp identified by this report will make it difficult for any bootcamp to achieve venture scale returns. This does not mean that short-term education is not a worthwhile endeavor or a viable business idea. Affordable alternatives to four year degrees are good and needed, and it’s theoretically possible for bootcamps to turn a profit while providing high quality education. However, all the available information suggests it’s difficult for bootcamps to be both profitable and to provide high quality education while growing at a rapid pace. It’s difficult to imagine a scenario where a coding bootcamp is a good venture capital investment.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;One approach for scaling exercised by a lot of coding bootcamps is &lt;a href=&quot;https://www.forbes.com/sites/oliversmith/2017/10/25/this-coding-school-has-a-business-model-to-boost-diversity-in-tech/#b89948e431e4&quot;&gt;focusing on diversity&lt;/a&gt; and promoting (and often subsidising) female and minority participation in order to boost their ratio in the tech workforce. The events unfolding in 2020 have &lt;a href=&quot;https://www.coursereport.com/blog/diversity-in-coding-bootcamps-report-2020&quot;&gt;accelerated this notion radically&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;According to &lt;a href=&quot;https://insights.stackoverflow.com/survey/2019&quot;&gt;StackOverflow&apos;s Developer Survey Results 2019&lt;/a&gt;, about 15.4% of all respondents have participated in a full-time developer program or training bootcamp. It should be noted, that bootcamps &lt;a href=&quot;https://insights.dice.com/2018/03/06/bootcamp-graduation-better-tech-salary/&quot;&gt;are not necessarily meant to be starting points for careers in tech&lt;/a&gt;, they represent an opportunity for reskilling and moving away from commoditised coding languages or unpopular fields.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_coding_bootcamp.png&quot; alt=&quot;Survey chart of the ways developers educate themselves, led by teaching yourself a new language, framework or tool at 85.5 percent&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Y Combinator intends to play a significant part in the future of non-traditional education, and has put their money where their mouth is attempting &lt;a href=&quot;https://medium.com/swlh/y-combinator-not-lambda-school-is-unbundling-education-bd6fdf0c78d7&quot;&gt;several promising tweaks to the business model&lt;/a&gt;... &lt;a href=&quot;https://web.archive.org/web/20200728090606/http://lloydalexander.me:80/can-lambda-school-become-a-100m-business-a-growth-case-study/&quot;&gt;without any major success yet&lt;/a&gt; though. It seems like even their bootcamp called Modern Labor tried &lt;a href=&quot;https://www.vice.com/en_us/article/yw878x/modern-labor-coding-bootcamp-will-pay-you-to-learn-to-code&quot;&gt;experimenting with connecting their educational business model with private sector projects through an ISA&lt;/a&gt;, but they had to realise that scaling education while keeping up the quality might be a more difficult task than anticipated. This leads us to our next point.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Seeing Silicon Valley&apos;s tendency to abuse their users in exchange for rapid company growth and in order to please their shareholders, the limits of scalability are both a challenge for coding bootcamp organisers and a helpful indicator for people searching for the right school. Reinventing the wheel did not work out for any school yet, but this does not mean that it won&apos;t in the future - you just need to make sure you are clear about the &lt;a href=&quot;https://qz.com/1038508/dev-bootcamps-president-explains-why-the-coding-bootcamp-shut-down/&quot;&gt;risks involved&lt;/a&gt;. &lt;/p&gt;
&lt;h3 &gt;2020&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;I am not sure how this year is going to go down in history, but I am confident that we all will have some anecdotes to tell about it. Observing how the lockdown, the social distancing measures, the pandemic-caused rise in unemployment and the WFH trend are accelerating the overdue seismic shift in higher education, the way we approach education as a means of personal socioeconomic improvement changes significantly.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/ad80aRfPZLU?si=pRju6uqsYVRTH_Sn&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

  &lt;p class=&quot;&quot; &gt;Sidenote: &lt;a href=&quot;https://www.classcentral.com/report/mooc-stats-pandemic/&quot;&gt;the popularity of MOOCs shot through the roof&lt;/a&gt; &lt;a href=&quot;https://www.classcentral.com/report/moocwatch-23-moocs-back-in-the-spotlight/&quot;&gt;during lockdown&lt;/a&gt;, but a certain level of sobering is already visible on the horizon as most learners have realised that they are just not motivated and/or disciplined enough to complete the courses they&apos;ve signed up for - the 2021 numbers will tell the full story. &lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_bootcamp-oas9yt.png&quot; alt=&quot;Newly registered learners in 2019 and 2020: Coursera 8M then 20M, edX 5M then 8M, FutureLearn 1.3M then 4M, Class Central 350k then 700k&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Critique and scandals &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;There are several &lt;a href=&quot;https://www.sfgate.com/tech/article/coding-bootcamps-are-they-worth-it-14341655.php&quot;&gt;solid articles&lt;/a&gt; for everyone looking for some research-backed advice on the upsides and downsides of attending a bootcamp, and various &lt;a href=&quot;https://medium.com/@ipresilia/i-went-to-a-coding-boot-camp-got-really-sick-and-failed-a72eed5998e2&quot;&gt;much more personal pieces&lt;/a&gt; from former students who suffered serious physical and mental stress during attending an intense school.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;While there were several small-to-medium scale scandals and closures around coding bootcamps in the last couple of years, they usually repeat the same mantra:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The bootcamp graduates are not fit to become software engineers at the highest-profile tech companies like Google,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Coding bootcamps use misleading marketing claims and post false outcome stats in terms of placement and salaries,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The quality of education participants receive is subpar and not worth the tuition fee (combine this with critique about ISAs as a financial instrument),&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The bootcamps are intense and stressful, some are ill-equipped to keep up with the fast pace.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Now, there are real life examples to all of the above claims (and usually they come in a bundle), but it shall be noted that the same &lt;a href=&quot;https://diff.substack.com/p/foxes-hedgehogs-vc-alpha-politics&quot;&gt;critique applies to for-profit higher ed in general&lt;/a&gt;. Universities and colleges (while partially being regulated) &lt;a href=&quot;https://twitter.com/stucchio/status/1230510524984524801&quot;&gt;don&apos;t receive scrutiny for unfulfilled outcomes&lt;/a&gt;, but the related financial debt has become one of the key economical issues of todays generation in the US, and therefore a highly politicized topic.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;One additional theme that comes back in a cyclical manner is related to VC investment in larger bootcamps that cannot deliver on their inflated market valuations. The &lt;a href=&quot;https://nymag.com/intelligencer/2020/02/lambda-schools-job-placement-rate-is-lower-than-claimed.html&quot;&gt;latest large scandal&lt;/a&gt; being about &lt;a href=&quot;https://www.theverge.com/2020/2/11/21131848/lambda-school-coding-bootcamp-isa-tuition-cost-free&quot;&gt;Y Combinator backed Lambda school&lt;/a&gt;, and its founder Austen Allred. An other sizeable scandal that made it into mainstream media was &lt;a href=&quot;https://www.buzzfeednews.com/article/daveyalba/datacamp-sexual-harassment-metoo-tech-startup&quot;&gt;a sexual misconduct story at DataCamp&lt;/a&gt;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The unquestionable gap in the education market filled by bootcamps and its importance in the future &lt;a href=&quot;https://www.politico.com/agenda/story/2019/01/16/coding-bootcamps-job-training-000869&quot;&gt;has been noticed by Washington as well&lt;/a&gt;, and we&apos;ll be witnessing even &lt;a href=&quot;https://www.politico.com/agenda/story/2019/01/16/coding-bootcamps-job-training-000869&quot;&gt;more political involvement in the coming years mostly focusing on regulation of financing and quality control&lt;/a&gt; - or even tying the two together by connecting institutional compensation with positive outcomes.&lt;/p&gt;
&lt;h3 &gt;Future outlook&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;The socioeconomic trends driven by digitalisation outside of Silicon Valley and the &quot;traditional&quot; tech sector will result in an even higher demand for skilled tech workers. This is going to &lt;a href=&quot;https://trainingindustry.com/articles/it-and-technical-training/the-future-of-coding-bootcamps/&quot;&gt;drive more experimentation&lt;/a&gt;, but also more consolidation to the (coding) bootcamp market. Business models combining online and offline learning, individual and corporate training and various forms of financing opportunities will to come and go. And politics is going to get involved with the notion to ensure outcomes for participants, regulation and comparisons with formal educations are already on the horizon.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://careerkarma.com/blog/bootcamp-market-report-2020/&quot;&gt;Differentiation and expansion into new non-tech areas should be expected&lt;/a&gt;, just as vertical expansion of the value chain turning educational institutions into end-to-end players on the tech (labour) market.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The future of the bootcamp business will be shaped by the general progress in edtech, and the growing, unsustainable socioeconomic bubble of US higher education, where &lt;a href=&quot;https://alexdanco.com/2019/12/17/ten-predictions-for-the-2020s/&quot;&gt;multiple&lt;/a&gt; analysts expect major disruption in the upcoming decade. The likely new entrants to the educational sector are &lt;a href=&quot;https://www.msn.com/en-us/news/technology/apple-is-offering-teachers-a-free-coding-course/ar-BB16xKik&quot;&gt;Apple&lt;/a&gt;, &lt;a href=&quot;https://techcrunch.com/2020/07/13/google-makes-education-push-in-india/&quot;&gt;Google&lt;/a&gt;, &lt;a href=&quot;https://web.archive.org/web/20201024110334/https://u2b.com/2020/07/06/microsoft-and-linkedin-offer-14-free-learning-paths-for-highly-in-demand-careers/&quot;&gt;Microsoft&lt;/a&gt; and Amazon: all of them have access to the capital and the tech competence to get involved in higher ed in a big way. Offering edtech tools and training, hardware to equip classrooms or enable learning from home, digital learning platforms combined with subscription models that end with a certification… you name it. Partnering with ivy league institutions who have raised their margins in an unprecedented manner in the last five decades will enable a more accessible college experience which however will be trimmed to deliver stakeholder value, &lt;a href=&quot;https://www.profgalloway.com/post-corona-higher-ed&quot;&gt;not one that is aimed at serving society across the board&lt;/a&gt;. As for coding bootcamps, most of them don’t have the brand recognition that would justify an investment from the largest players, but the opportunity to partner up with medium-sized or local tech players is still huge.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Before we close this chapter, let&apos;s reflect on what the key learnings of the last decade really are. Here is some more &lt;a href=&quot;https://www.fastcompany.com/40471956/this-is-what-coding-bootcamps-need-to-do-to-beat-the-backlash&quot;&gt;operative advice for coding bootcamps&lt;/a&gt; by Jonathan Lau, the founder of &lt;a href=&quot;http://switchup.org&quot;&gt;SwitchUp.org&lt;/a&gt;, one of the two leading review site for coding bootcamps:&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“DON’T SCALE UP AT THE COST OF QUALITY  &lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;REFLECT THE NEEDS OF LOCAL MARKETS&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;KEEP YOUR OFFERINGS CURRENT AND RELEVANT&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;COMPETE ON CREDIBILITY ABOVE ALL ELSE&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;EMBRACE FEEDBACK FROM THE MARKETPLACE”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;&quot; &gt;Similarly, Jason Moss (President and Founder of the data science bootcamp Metis) &lt;a href=&quot;https://www.thisismetis.com/blog/6-lessons-learned-in-6-years-at-metis&quot;&gt;has provided his POV&lt;/a&gt; on the last years working on turning his school to a success story: &lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“Embrace continuous transformation. Nothing is “right” for long.  &lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Think digital first. Plan for scale from the get-go.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Fish in large ponds.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Don’t make assumptions lightly. Be clear about big decisions.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;There is excellence in focus.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Invest in people and reputation. Good things will follow.”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 &gt;Closer&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;To finish off, allow me go back to &lt;a href=&quot;https://www.freecodecamp.org/news/coding-bootcamp-handbook/&quot;&gt;The Coding Bootcamp Handbook&lt;/a&gt; for a minute and point out a few words that are on my mind since the decision about founding Pipeline Academy, the first school for data engineering has been made:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;“[...] new coding bootcamps have something to prove. Their teachers and staff will work like crazy to ensure the school succeeds. They&apos;ll try their hardest to train you. They&apos;ll help you get a good job so they can get a win under their belt and onto their testimonials page.”&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Whatever the future vision will look like (and trust me, I already have a not-so-vague idea about it), this will be the guiding light when starting the journey. And this is what you should expect from us. Nothing less.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - July 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/08/13/the-data-janitor-letters-july-2020.html"/>
   <updated>2020-08-13T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/08/13/the-data-janitor-letters-july-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://hakibenita.com/sql-tricks-application-dba&quot;&gt;Some SQL Tricks of an Application DBA&lt;/a&gt;&lt;br&gt;&lt;em&gt;Haki Benita, Development Team Lead, PCENTRA&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Databases are the backbone of most modern systems, so taking some time to understand how they work is a good investment for any developer!&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20201002001006/https://counting.substack.com/p/sessions-for-analysis-the-eternal&quot; target=&quot;_blank&quot;&gt;Sessions for analysis, the eternal fiction&lt;/a&gt;&lt;br&gt;&lt;em&gt;Randy Au, Quantitative UX Researcher, Google &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;There are a vast multitude of hypotheses and potential narratives we could attach to any session.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/omio-engineering/our-journey-to-a-new-data-warehouse-fc4488e46761&quot; target=&quot;_blank&quot;&gt;Our journey to a new data warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ana Gulevskaia, BI Analyst, Omio&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;How do the benefits of 3NF and SF fit in here?&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.marksblogg.com/omnisci-macos-macbookpro-mbp.html&quot; target=&quot;_blank&quot;&gt;1.1 Billion Taxi Rides using OmniSciDB and a MacBook Pro&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/marklit82&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Mark Litwintschik&lt;/em&gt;&lt;/a&gt;&lt;em&gt; #BigData Consultant&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The Q1 time is the fastest for any workstation benchmark I&apos;ve done. To get this level of performance on a regular piece of office equipment is a big game changer.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
Big Data Small GPU, No Problem&lt;br&gt;&lt;em&gt;Rodrigo Aramburu, CEO, BlazingSQL&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;BlazingSQL is no longer limited by available GPU memory for query execution.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@paigeonthewing/evolution-of-the-modern-data-warehouse-f8b8d616149d&quot; target=&quot;_blank&quot;&gt;Evolution of the Modern Data Warehouse&lt;/a&gt;&lt;br&gt;&lt;em&gt;Paige Roberts, Open Source Relations Manager, Vertica&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A data warehouse is essentially a business-driven, enterprise-centric and technology-based solution.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.thoughtworks.com/insights/blog/how-much-can-you-trust-your-data&quot; target=&quot;_blank&quot;&gt;How much can you trust your data?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ellen König, Senior Data Engineer, ThoughtWorks&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data quality assessments are an effective, but often overlooked way to make your company’s data products more trustworthy.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://artem.krylysov.com/blog/2020/07/28/lets-build-a-full-text-search-engine/&quot; target=&quot;_blank&quot;&gt;Let&apos;s build a Full-Text Search engine&lt;/a&gt;&lt;br&gt;&lt;em&gt;Artem Krylysov, Senior Software Engineer, Datadog&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Despite its simplicity, it can be a solid foundation for more advanced projects.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://maxkostinevich.com/blog/serverless-geolocation/&quot; target=&quot;_blank&quot;&gt;How to make simple Geolocation service&lt;/a&gt;&lt;br&gt;&lt;em&gt;Max Kostinevich&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;On Cloudflare Workers I&apos;ve got almost x10 better performance in comparison to AWS.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;&lt;a href=&quot;https://arrow.apache.org/blog/2020/07/24/1.0.0-release/&quot; target=&quot;_blank&quot;&gt;Apache Arrow 1.0.0 Release&lt;/a&gt;&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The 1.0.0 release indicates that the Arrow columnar format is declared stable, with forward and backward compatibility guarantees.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>I'm not an expert. I have experience.</title>
   <link href="https://dataengineering.academy/2020/08/07/data-engineering-data-janitor-keynotes.html"/>
   <updated>2020-08-07T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/08/07/data-engineering-data-janitor-keynotes.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;I don&apos;t like to talk. I don&apos;t like to talk at conferences. I&apos;m an introvert at heart. Still, sometimes one has to step out of the comfort zone and stand for things that deserve to be stood for. All my talks were born this way, some more technical, some more reflective on the practicalities of data life.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;In 2016, at &lt;a href=&quot;https://berlinbuzzwords.de/&quot;&gt;Berlin Buzzwords&lt;/a&gt; in Berlin I shared our journey of how we migrated a data stack from AWS to Azure. How we managed to do it with open sourcing tools as roadkills and what we’ve learnt about the barebone necessities ending up within the sole body of a Raspberry Pi may shock the absolute believers of distributed computing. An adventure in bash, make and SQL with a detour in Moore&apos;s law and falling memory prices. You can find the slides &lt;a href=&quot;https://www.slideshare.net/soobrosa&quot;&gt;right here&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/QhXPANTd9nE?si=ELRFoHryl_chNRYD&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/oKWmg3oBJgc?si=6i9aQbLV9EDejhJk&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/LTJNnlBBzuw?si=9ix7eVA5ZW-wHQxA&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - June 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/07/20/the-data-janitor-letters-june-2020.html"/>
   <updated>2020-07-20T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/07/20/the-data-janitor-letters-june-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20200811155850/https://better.engineering/2020-06-24-wizard-part-ii/&quot; target=&quot;_blank&quot;&gt;I gave the business what they asked for and they never used it&lt;/a&gt;&lt;br&gt;&lt;em&gt;Kenny Ning, Data Engineer, Better&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;An addition to the long list of unused ML projects: do the simplest solution first, modeling is never done, failure is okay.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://veekaybee.github.io/2020/06/09/ml-in-prod/&quot;&gt;Getting machine learning to production&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/vboykis&quot;&gt;Vicki Boykis&lt;/a&gt;&lt;em&gt;, Machine Learning Engineer, automattiC&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Deploying is hard. Deep learning is deceptively easy. Go for prebuilt as much as possible. Understand networking and scale. Iterate quickly.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.piekniewski.info/2020/06/08/ai-the-no-bullshit-approach/&quot; target=&quot;_blank&quot;&gt;AI – the no bullshit approach&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/filippie509&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Filip Piekniewski&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Scientist, Accel Robotics&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Common sense will emerge only when a connectionist like system will have a chance to develop the internal symbols to represent the relationships in physical world.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://huyenchip.com/2020/06/22/mlops.html&quot; target=&quot;_blank&quot;&gt;What I learned from looking at 200 machine learning tools&lt;/a&gt;&lt;br&gt;&lt;em&gt;Chip Huyen, Engineer, Snorkel AI&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If you have to choose between engineering and ML, choose engineering. It’s easier for great engineers to pick up ML knowledge, but it’s a lot harder for ML experts to become great engineers.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.aleksandra.codes/tech-content-consumer&quot; target=&quot;_blank&quot;&gt;Most tech content is bullshit&lt;/a&gt;&lt;br&gt;&lt;em&gt;Aleksandra Sikora, Software Engineer, Hasura&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Realize that there&apos;s tons of misconception in the world. Adapt solutions to your particular use case. Your solutions are not any worse than the ones on the internet.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/omio-engineering/how-we-migrated-our-data-warehouse-from-redshift-to-bigquery-89f33988e1b5&quot; target=&quot;_blank&quot;&gt;How we migrated our data warehouse from Redshift to BigQuery&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/rahulj51&quot; target=&quot;_blank&quot;&gt;&lt;em&gt;Rahul Jain&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Principal Engineering Manager, Data engineering and BI platform, Omio/GoEuro&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A journey step-by-step.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
In search of speed — debugging Elasticsearch performance&lt;br&gt;&lt;em&gt;Martin Iotchev, Fullstack Software Engineer, Prezi&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The underlying hardware plays a significant role in the performance of an Elasticsearch cluster. Provisioning larger data nodes will yield better performance as compared to the smaller default nodes currently used in production. Furthermore, a cluster with more shards will perform better on larger data sets.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://lwn.net/Articles/822568/&quot; target=&quot;_blank&quot;&gt;Lightweight alternatives to Google Analytics&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ben Hoyt, Software Engineer, Compass&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;For site owners who just need basic traffic numbers, GoatCounter and Plausible both seem like excellent options. Those who like more visual polish and documentation might prefer Plausible; those who value a more developer-friendly tool with easy self-hosting will probably prefer GoatCounter.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>This is not a test, this is a summer camp - part #2</title>
   <link href="https://dataengineering.academy/2020/07/08/this-is-not-a-test-this-is-a-summer-camp-part-2.html"/>
   <updated>2020-07-08T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/07/08/this-is-not-a-test-this-is-a-summer-camp-part-2.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;This is the second part of the story about setting up a live remote coding workshop in the midst of an economic downturn. Part one is about the circumstances that led to all this, you can &lt;a href=&quot;/2020/05/26/this-is-not-a-test-this-is-a-summer-camp-part-1.html&quot; target=&quot;_blank&quot;&gt;read it here&lt;/a&gt;. Part two is about how we&apos;ve rocked the workshop itself and about the feedback we&apos;ve received after.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;When I started reading the responses from the eight participants of the &lt;a href=&quot;/data-engineering-summer-camp/&quot;&gt;Data Engineering Summer Camp&lt;/a&gt; to the feedback questionnaire I&apos;ve shared with them right after the workshop, the feeling this whole situation sparked in me reminded me of the effect named after the 1950 Jidaigeki crime movie &lt;a href=&quot;https://en.wikipedia.org/wiki/Rashomon&quot;&gt;Rashomon&lt;/a&gt; by Akira Kurosawa. Karl G. Heider used the term &lt;a href=&quot;https://en.wikipedia.org/wiki/Rashomon_effect&quot;&gt;&quot;The Rashomon Principle&quot;&lt;/a&gt; to refer to the effect of the subjectivity of perception on recollection, by which observers of an event are able to produce substantially different but equally plausible accounts of it. Even though the ten of us (including Daniel and myself) were not external observers of the Summer Camp but active participants of the event, when putting these ten individual recollections next to each other the data becomes a physical manifestation of this particular gestalt.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/xCZ9TguVOIA?si=mxUiWswiVjqmB-OW&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;My point: whoever is reading this should keep in mind that this summary is inherently biased and there is no &lt;em&gt;one&lt;/em&gt; correct interpretation of the event in question. What I am looking for are shared patterns in the multitude of witness perspectives involved that indicate a direction for the path to our own improvement.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;But then you see a tweet like this from one of our happy students and you instantly forget about the semi-scientific approach you were about to force upon yourself:&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/flavioclesio.png&quot; alt=&quot;Flávio Clésio&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Source: &lt;a href=&quot;https://twitter.com/flavioclesio&quot;&gt;Flávio Clésio&apos;s twitter&lt;/a&gt; (&lt;a href=&quot;https://web.archive.org/web/20200611124746/https://twitter.com/flavioclesio/status/1271059360320491520&quot;&gt;saved version&lt;/a&gt;), &lt;a href=&quot;https://twitter.com/soobrosa?lang=en&quot;&gt;@soobrosa&lt;/a&gt; is Daniel&apos;s twitter handle.&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt; First, let&apos;s go back to where we left off the last time.&lt;/p&gt;
&lt;h3 &gt;The Data Engineering Summer Camp&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Quick recap: between the 18-22nd May (Monday-Friday) we&apos;ve held the first Summer Camp by Pipeline Academy with eight participants in the virtual classroom. &lt;strong&gt;The purpose was supporting people and businesses by teaching them a valuable competence through a hands-on project and learning about our own areas of improvement through the process.&lt;/strong&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The agenda was designed to cover the most important building blocks of setting up a simple data product: an introduction to data engineering, ETL basics, some SQL and deployment. Since we prefer &lt;em&gt;actual building&lt;/em&gt; to just &lt;a href=&quot;/2020/05/04/its-time-to-build-data-pipelines.html&quot;&gt;talking about building&lt;/a&gt;, the decision to focus on real implementation was quickly made. The below schedule reflects this attempt with 2.5 days of classroom experience aka &apos;campsite&apos; (frontal teaching, live coding, group discussions, presentations and demos, Q&amp;amp;A) and 2.5 days of individual or team adventures (in duos) for implementing the data pipeline.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/weekly_agenda.png&quot; WIDTH=800 alt=&quot;The weekly agenda: a Monday to Friday timetable of sessions with lunch breaks&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;The week before the workshop the participants received homework. At first glance, the task was fairly simple and straightforward, but it included some challenges that could be solved in various ways, depending on what tools and methods a student prefers. Imagine you&apos;re asked to travel from point A to point B within a city: there are plenty of alternative ways of completing the task picking different routes, means of transportation, and the chosen way serves as a testament to the preferences and skillset of the traveler. We used this dry run to get a better understanding of what our cohort is comfortable with vs. what topics we need to put more emphasis on to achieve the desired outcome for the week.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The first morning was about getting familiar with each other, some chit-chat about our daily routines during corona and our different experiences with working with data. We&apos;ve moved on to discussing what data engineering is and how the skills that make a data engineer relate to the participants&apos; individual skillset. Here&apos;s an example: a data scientist usually has more advanced mathematical know-how than a frontend engineer, but the latter is likely to have seen or written more maintainable code (especially collaboratively in larger teams). The afternoon was spent with an overview of ETL procedures (&lt;a href=&quot;https://en.wikipedia.org/wiki/Extract,_transform,_load&quot;&gt;extract, transform, load&lt;/a&gt;).&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The first day was very exciting but exhausting as well. I was relieved that we&apos;ve managed to capture the attention of the participants as they kept coming back the following days, without any measurable churn rate. Some of the students had to skip a couple of hours of classes during the week due to bureaucratic obligations and work emergencies, and sometimes people did not join the class instantly due to issues with their internet provider at home (&lt;a href=&quot;https://www.tagesspiegel.de/wirtschaft/internet-berliner-surfen-besonders-langsam/21071356.html#:~:text=Internet%20Berliner%20surfen%20besonders%20langsam,Internetverbindungen%20unter%20den%20deutschen%20Gro%C3%9Fst%C3%A4dten.&amp;amp;text=Berlin%20liegt%20auf%20Rang%2078%20mit%20einer%20Geschwindigkeit%20von%20knapp,Einwohner%20der%20anderen%20deutschen%20Gro%C3%9Fst%C3%A4dte.)&quot;&gt;Berlin, du bist so wunderbar&lt;/a&gt;). We&apos;ve had one single person who pretty much gave up on delivering his solution (albeit staying on board until the end) as a result of unforeseen duties that hijacked their time and attention for the week.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/daniel_the_data_janitor.png&quot; WIDTH=800 alt=&quot;Daniel Molnar, the Data Janitor&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;This is not an illustration.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Part of the plan was letting the hands of the students go so they can explore the newly learned methodologies and tools by themselves and combine it with their existing knowledge. &lt;strong&gt;Our aim was to enable independent work and push for applied creative problem-solving&lt;/strong&gt;, while staying available for everyone for questions and support via chat. The student feedback confirmed our assumption: this was a highly productive segment of the week that accelerated the pace of learning rapidly.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;It was remarkable how much interest there was in seemingly niche data engineering topics and obscure tools: some people were enthusiastic and some were a bit more skeptical about building a data pipeline with some unfamiliar puzzle pieces they were not accustomed to (i.e. SQLite), but most students were open to exploration and experimentation. Five days passed by in an instant.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;As a closing event, the participants who have worked in teams had to present their data products on Friday. The solutions showed two examples of highly successful and productive collaborations, and two teams had hit roadblocks that could be clearly identified and discussed to pave the way for a late delivery. In addition to the course certificate, Daniel and I decided to gift Udemy courses to the students in order to make sure that they continue pushing themselves forward on the endless path of data engineering.&lt;/p&gt;
&lt;h3 &gt;Feedback time&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;We&apos;ve asked the students to fill out a survey so we get a better understanding of how they see the Summer Camp, you can read some of their sentiments in &lt;em&gt;italic&lt;/em&gt; below. The learnings were many:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;General concept: there is a large demand for data engineering skills on the market, and even for more seasoned veterans, finding the right resources to access and learn this know-how in a structured and methodologically proven way is almost impossible. We&apos;re on the right track.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p class=&quot;&quot; &gt;The quality of the content and the delivery was rated very high across the board and our students had fun. Oh, you are asking about an NPS score you crazy marketing researcher: it&apos;s excellent, baby. Good stuff. &lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;Overall, very satisfied with the course, Daniel is a great teacher and this course was very valuable to me.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;In-person interaction in a classroom vs. live remote learning via Zoom is something polarising, however the majority of students prefers personal contact. Reasons: higher productivity, more interactions with peers (aka camaraderie) and online education is just not as direct and immersive. At the same time, the remote live experience is the right way to go when it comes to a fallback plan in case COVID-19 is coming back for seconds in the winter.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p class=&quot;&quot; &gt;Teamwork makes the dream work: we&apos;re strong believers of the value that lies in working in teams, and this is something we&apos;ll keep as integral part of our methodology. &lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;The team adventures helped me to see other ways to solve the same problem and I think this should be kept for the next courses for sure.&quot;, &quot;They helped consolidate the learnings by rolling up your sleeves and solving challenges practically with fellow peers or by yourself.&quot;, &quot;Yes, got unblocked a few times by my partner.&quot;&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p class=&quot;&quot; &gt;We need to focus more on defining clear expectations for the challenges and tasks: striking a balance between allowing room for creative problem-solving and defining clear outcomes is the name of the game for engineering exercises. This is something we&apos;re going to fine tune &lt;a href=&quot;/apply/&quot;&gt;by the launch of our first cohort in the fall&lt;/a&gt;. &lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;&lt;li&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;I was demotivated at times. The instructions were too vague and I got lost in useless details (flask issue on pythonanywhere).&quot;&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Assessing the difficulty of the course based on our small sample is tough: depending on the student&apos;s level of experience, general mindset, life-circumstances outside of the summer camp and approach to self-assessment, students end up with a variety of sentiments. Fortunately a proper bootcamp experience with 10+ weeks of full-time education allows for more personalised mentoring.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/data_engineering_checklist.png&quot; alt=&quot;Checklist: don&apos;t write code, solve the problem; keep it simple stupid; separation of concerns; delete, remove, retire; code is dependency; others&apos; code is dependency squared&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;More fluff&lt;/h3&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;The course delivered a good blend between technical hands-on and practical aspects about the Data Engineering profession and roles. The course has a modern approach to combine the pragmatism of the data activities (e.g. reliability, well thought changes) with modern approaches (e.g. online deployment, orchestration, monitoring, etc).&quot;&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Curricula and roadmaps for getting started with data science are plenty, data engineering however requires a different type of didactics and a learning environment that mimics the engineering and collaboration processes at companies leveraging technology.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Adam, long time data science teacher has written up his experiences at the summer camp &lt;a href=&quot;https://web.archive.org/web/20200715072622/https://adgefficiency.com/pipeline-data-engineering-academy/&quot;&gt;on his blog&lt;/a&gt; focusing on how this week&apos;s learnings enabled him to move forward with his own projects. Make sure to &lt;a href=&quot;https://web.archive.org/web/20200715072622/https://adgefficiency.com/pipeline-data-engineering-academy/&quot;&gt;check out his post&lt;/a&gt; for a more technical POV (full disclosure: I&apos;ve had the pleasure of working with Adam for a while and consider him a friend):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;&quot;As someone who has taught data science for a while, the most impressive thing was the simplicity of the stack. It is very easy when teaching to complicate things for students, or to teach complex tools that confuse more than help.&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;It can’t be understated the power of leaving after five days with a working product. The value of being able to see and interact with your data is huge, for spotting problems with your data pipeline to showing off to customers (or employers!).&quot;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;The Summer Camp was my major highlight of the lockdown: it turned out to be a productive validation for Pipeline Academy, and furthermore we&apos;ve managed to support people and share valuable knowledge while staying true to our principles - transparency, collaboration, pragmatism and common sense.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I hope to see you at our campus in the fall.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>5 + 1 Keynotes to remember</title>
   <link href="https://dataengineering.academy/2020/06/21/data-engineering-keynotes-to-remember.html"/>
   <updated>2020-06-21T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/06/21/data-engineering-keynotes-to-remember.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Unfortunately &lt;a href=&quot;https://www.aleksandra.codes/tech-content-consumer&quot;&gt;most tech content is bullshit&lt;/a&gt;, still there are solid cornerstones of the engineering society. They do raise their voices and did keynotes at conferences. These are the keynotes I watched more than once, some are dire, some are hilarious. Use them for inspiration, learning or simply just for fun.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://www.youtube.com/watch?v=pW-SOdj4Kkk&quot;&gt;Jonathan Blow: Preventing the Collapse of Civilization&lt;/a&gt; at DevGAMM 2019&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;He is an American video game designer and programmer, who is best known as the creator of the independent video games Braid and The Witness.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/pW-SOdj4Kkk?si=_M1KdVmztfTunp_i&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://vimeo.com/95066828&quot;&gt;James Mickens: Computers are a Sadness, I am the Cure&lt;/a&gt; at Monitorama PDX 2014&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Caveat: I do maintain the &lt;a href=&quot;http://soobrosa.info/static/2016-04-05_my-humble-james-mickens-shrine-aka-the-only-real-combined-cs-degree-and-mba-you-will-ever-need.html&quot;&gt;James Mickens Shrine&lt;/a&gt; so it was a hard decision. The classic.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe src=&quot;https://player.vimeo.com/video/95066828?h=2b5433dc49&amp;portrait=0&quot; width=&quot;640&quot; height=&quot;360&quot; frameborder=&quot;0&quot; allow=&quot;autoplay; fullscreen; picture-in-picture&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.youtube.com/watch?v=RZ4Sn-Y7AP8&quot;&gt;David Beazley: Discovering Python&lt;/a&gt; at PyCon 2014&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;What happens when you lock a Python programmer in a secret vault containing 1.5 TBytes of C++ source code and no internet connection? Python as a secret weapon of &quot;discovery&quot; in an epic legal battle.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/RZ4Sn-Y7AP8?si=g1joMdah8I9pzOrk&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.destroyallsoftware.com/talks/the-birth-and-death-of-javascript&quot;&gt;Gary Bernhardt: The Birth &amp;amp; Death of JavaScript&lt;/a&gt; at PyCon 2014&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;This science fiction / comedy / absurdist / completely serious talk traces the history of JavaScript, and programming in general, from 1995 until 2035&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/the-birth-and-death-of-javascript.png&quot; alt=&quot;Video still from the talk “The Birth &amp; Death of JavaScript”&quot;&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.youtube.com/watch?v=lKXe3HUG2l4&quot;&gt;Joe Armstrong: The Mess We&apos;re In&lt;/a&gt; at Strange Loop 2014&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The late Joe Armstrong was one of the inventors of Erlang and a great storyteller.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/lKXe3HUG2l4?si=4yi93vKDcohG3sGJ&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;+1: &lt;a href=&quot;https://www.youtube.com/watch?v=EK4qctJOMaU&quot;&gt;Chris Ford: African Polyphony &amp;amp; Polyrhythm&lt;/a&gt; at Strange Loop 2016&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Now this is an odd choice, isn&apos;t it? It&apos;s about models, reality, data, interpretation, culture. Also in-browser music.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/EK4qctJOMaU?si=ewv_WmibvkgovI3-&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - May 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/06/12/the-data-janitor-letters-may-2020.html"/>
   <updated>2020-06-12T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/06/12/the-data-janitor-letters-may-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://saagarjha.com/blog/2020/05/10/why-we-at-famous-company-switched-to-hyped-technology/&quot;&gt;Why we at $FAMOUS_COMPANY Switched to $HYPED_TECHNOLOGY&lt;/a&gt;&lt;br&gt;&lt;em&gt;Saagar Jha&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Today we are making some of the code that we can afford to open source available on our GitHub page. It is useless by itself and is heavily tied to our infrastructure, but you can star it to make us seem more relevant.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://mebassett.info/ai-useless-for-business&quot;&gt;Why is Artificial Intelligence So Useless for Business?&lt;/a&gt;&lt;br&gt;&lt;em&gt;Matthew Eric Bassett&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We face two more immediate problems: a lack of data and a lack of awareness.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://techcrunch.com/2020/05/06/no-cookie-consent-walls-and-no-scrolling-isnt-consent-says-eu-data-protection-body/&quot;&gt;No cookie consent walls — and no, scrolling isn’t consent, says EU data protection body&lt;/a&gt;&lt;br&gt;&lt;em&gt;Natasha Lomas, TechCrunch&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;There’s increasing pressure on regulators to actually enforce the rules — with GDPR’s two year anniversary fast approaching.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20200514164445/https://towardsdatascience.com/dont-become-a-data-scientist-ee4769899025&quot;&gt;Don’t Become a Data Scientist&lt;/a&gt;&lt;br&gt;&lt;em&gt;Chris, Towards Data Science&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The advice I give when someone asks me how to get into data science. Become a software engineer instead.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.theregister.co.uk/2020/05/13/teradata_prospects_new_ceo/&quot;&gt;Teradata switches CEOs mid-flight while being eaten alive in the cloud, but it&apos;s not game over yet for data warehouser&lt;/a&gt;&lt;br&gt;&lt;em&gt;Lindsay Clark, The Register&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Last year, Gartner noted that just two vendors – AWS with Redshift and Microsoft with Azure SQL Data Warehouse – accounted for 70 per cent of the revenue growth in the overall DBMS market for 2017 and 2018.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.timescale.com/blog/multi-node-petabyte-scale-time-series-database-postgresql-free-tsdb/&quot;&gt;TimescaleDB&lt;/a&gt;&lt;br&gt;&lt;em&gt;Ajay Kulkarni, CEO and co-founder, Timescale&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A multi-node, elastic, petabyte scale, time-series database on Postgres for free.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.youtube.com/watch?v=fGG9dApIhDU&quot;&gt;Carnegie Mellon University Database Group - Introducing ClickHouse&lt;/a&gt; (VIDEO)&lt;br&gt;&lt;em&gt;Robert Hodges, CEO, Altinity &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The Fastest Data Warehouse You&apos;ve Never Heard Of&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;&lt;a href=&quot;https://news.ycombinator.com/item?id=23206566&quot;&gt;What every software engineer should know about Apache Kafka | Hacker News&lt;/a&gt;&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;&quot;Do you happen to know what sort of thinking leads a team down this path? It seems a fundamental mistake. Resume driven development?&quot;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>This is not a test, this is a summer camp - part #1</title>
   <link href="https://dataengineering.academy/2020/05/26/this-is-not-a-test-this-is-a-summer-camp-part-1.html"/>
   <updated>2020-05-26T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/05/26/this-is-not-a-test-this-is-a-summer-camp-part-1.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;This is the first part of the story about setting up a live remote coding workshop in the midst an economic downturn. Part one is about the circumstances that led to all this and the way we&apos;re going with the flow.&lt;/p&gt;
&lt;h3 &gt;Let&apos;s try to do this differently &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;I knew the road to starting a company is rocky, but dealing with the sudden surge of a pandemic was not on my radar. Once the risks of COVID-19 became clearer and the measures to stop the spread of the virus started to take shape back in March, their impact on our social lives and the projected consequences for our economy forced all of us to start coping with the idea of a future that is way more uncertain than just a couple of weeks before. If this mental state reminds you of &lt;a href=&quot;https://hbr.org/2020/03/that-discomfort-youre-feeling-is-grief&quot;&gt;what grief must feel like&lt;/a&gt;, you&apos;re not wrong. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I guess Daniel and I have arrived at the final stage of the grieving process relatively quickly: acceptance for us meant, let&apos;s move ahead and change our initial plan of opening our own coding bootcamp for now and figure out how we can contribute in a meaningful way to help people. Unlike with other natural disasters, there is no blueprint for what you should do in order to support your community (beyond the social distancing and hygiene measures), so we had to get creative. That&apos;s how we ended up with the concept of a free live remote data engineering course for individuals and businesses who (we think) can leverage the value of this type of expertise to a) reduce the risk of unemployment by diversifying their competencies and to b) enable organisations to become more efficient and therefore more resilient to an economic slump. Teaching data engineering might not sound as compassionate as organising the umpty-umpth hackathon against corona where the only comprehensible output is the participants’ vehement virtue signalling under a new hashtag, but hey, that&apos;s what we think we&apos;re good at.  It was reassuring to see a lot of companies launching their makeshift initiatives to actually help others in need parallel to us in the following weeks, I think &lt;a href=&quot;https://web.archive.org/web/20200811031037/https://www.tomhayton.com/&quot;&gt;Tom Hayton put it very well in his new book&lt;/a&gt; comparing this process to the movie classic, Groundhog Day:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;[...] &lt;em&gt;Phil eventually realises that he can turn the situation to his advantage, mastering new skills and trying to win over a love interest. Only when he focuses his energy on serving others, though, does he achieve true fulfilment and happiness... [...]&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/GncQtURdcE4?si=1bsydsSw0aeMAC7J&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;Before you accuse us of being not just utterly handsome and smart, but extremely generous and altruistic as well, allow me to set the record straight: we have told all interested participants that ex post facto we would like to have one thing in exchange: their feedback.&lt;/p&gt;
&lt;h3 &gt;Preparation for validation&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;So the overarching concept of the Pipeline Data Engineering Summer Camp was born, but the details had to be worked out:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What is a good intersection of useful, valuable data engineering expertise and practical content that we can transfer to our participants in a five day workshop?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What are the technical and methodological constraints of using live videoconferencing instead of a real classroom experience? How do we need to change our approach?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What can we expect from our participants? What should they expect from us?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;What platforms and tools should we use for the virtual classroom and communication in general?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;How do we make it motivating and fun?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;How do we stay on top of our personal obligations (babysitting, homeschooling, watching Tiger King etc.) and manage WFH in parallel?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Don&apos;t get me wrong, it&apos;s not that there is no literature on these topics or we haven&apos;t thought about them before, but our approach based on our gut feeling has to be verified, adjusted and tested before launch.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;This workshop is supposed to resemble a miniature version of our original plan, a full-blown coding bootcamp for data engineering. Our vision for creating a broader platform for the data workers of the future remains untouched, we&apos;ve only adjusted the next step on the roadmap. This is important because in order to understand our rapidly growing niche even better we have to hear what works and what doesn&apos;t, especially now, when consumer sentiment and attitude are on the brink of a major realignment. Needless to say, we have a very good idea about what we should be teaching and who our ideal corporate partners would be, yet the question remains: what do our customers think about our ideas? We have to validate our assumptions and gain direct feedback from our target group(s), and this is what the Summer Camp can be used as a vehicle for.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Ultimately, we&apos;re trying to create a win-win situation here: first and foremost our participants should benefit from the workshop and based on the gained experience and feedback we should get better at what we do as well. Even though &lt;a href=&quot;https://www.theatlantic.com/health/archive/2011/11/does-altruism-actually-exist/248074/&quot;&gt;our initial motive came from a selfless place&lt;/a&gt;, in the end our intention has forked: helping others while helping ourselves.&lt;/p&gt;
&lt;h3 &gt;Less talk, more rock&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Carving out the details included sticking to our guiding principles when doing business: creating value for every stakeholder involved through an enjoyable high-quality experience. Achieving this required some ground rules, so Daniel and I agreed to the following:&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;No more than eight participants so we have sufficient time to focus on each person&apos;s individual challenges,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;As little frontal teaching as possible, individual problem solving and working in pairs is going to be encouraged,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The goal for everyone is to deliver an end-to-end data product in five days, so the workshop will be extremely hands-on,&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The scope and the way we interact with the participants should resemble our fairly unique approach to the subject matter (&quot;the little black dress vs. fast fashion&quot;),&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The party is going to be invite only, it&apos;s meant for the network of our networks in the first place to make sure that we have participants who benefit from the camp. &lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;After we&apos;ve set the date to end of May, we started reaching out to our respective networks in order to fill the virtual room with the right people. Our assumption in April was that with the lockdown still going strong and with online courses experiencing a renaissance (especially the ones revolving around data), with the job market being in a major turmoil and with organisations still trying to get accustomed to the new normal, the demand for &lt;a href=&quot;/data-engineering-summer-camp/&quot;&gt;a free workshop about a red-hot topic like data engineering&lt;/a&gt; will be through the roof.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Well, this turned out only to be half-right: we&apos;ve offered a spot to only 15-18 data leaders from Berlin (CTOs, CEOs, engineering leads, head of product etc.) asking them to nominate somebody from their respective data/tech teams for the workshop if they think they have a person who&apos;s a good fit (has the basics, is eager to learn and would benefit from the camp) and in case they can release them from work for a full week. In addition, we had two spots reserved for people who we knew directly and expressed interest in taking part. After two weeks with some back-and-forth about availability, we&apos;ve had eight people on the list.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;In the meantime, both of us were busy setting up the tools required for class communication (this involved some cursing), benchmarking competing services for live remote education (a lot more cursing), doing actual live rehearsals to get accustomed to the online format and making sure that the agenda and the schedule work fine. We&apos;ve picked Zoom for videoconferencing not because it is good in any way, but it&apos;s the best out there. Using it was just like you would expect it: like sitting on a plane that is being assembled during takeoff while the flight attendant explains the nuanced meaning of the word &lt;a href=&quot;https://status.zoom.us/&quot;&gt;&apos;operational&apos;&lt;/a&gt; and &lt;a href=&quot;https://www.forbes.com/sites/kateoflahertyuk/2020/05/08/zoom-security-you-need-to-know-about-these-3-new-features-arriving-tomorrow/#38471b5c60f8&quot;&gt;their unique concept of security&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;

&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/6J538b-OLRU?si=Va--FA6id0psXNmF&quot; title=&quot;YouTube video player&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; allowfullscreen&gt;&lt;/iframe&gt;

&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Ready for liftoff&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Just a couple of days before the workshop one participant had to drop off due to an unforeseen conflict in their schedule. By sheer coincidence, the same day a person reached out to Daniel via LinkedIn: he was looking for some guidance to get his foot in the door in the world of data engineering and prepare for a new chapter of his career. He was not aware of the summer camp as we kept it under the radar (no public promotion at all), but after he learned about the scope, he was in. We&apos;ve ended up with a data scientist, data engineers with different levels of experience, two data analysts and a software engineer with background in mobile and backend development on the roster. It was the perfect mix of varying skill levels and competencies, the fact that we had people from prestigious startups and scale-ups from the Berlin scene was just icing on the cake. We were ready to go.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;/2020/07/08/this-is-not-a-test-this-is-a-summer-camp-part-2.html&quot;&gt;&lt;em&gt;Part #2&lt;/em&gt;&lt;/a&gt;&lt;em&gt; is about how the summer camp went and the feedback we&apos;ve received. Follow us on social media and sign up to our newsletter to learn more about our adventure! &lt;/em&gt;&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>Own Your Data - like fo' real.</title>
   <link href="https://dataengineering.academy/2020/05/17/own-your-data-like-fo-real.html"/>
   <updated>2020-05-17T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/05/17/own-your-data-like-fo-real.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Have you ever been wondering why Facebook has your memories and why you don&apos;t have them?&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I had this urge last Friday CET 4:40 PM to try to get my own data from different digital services that handle it. As an EU citizen, this is my right according to the &lt;a href=&quot;https://eur-lex.europa.eu/eli/reg/2016/679/oj&quot;&gt;Art. 20 GDPR &quot;Right to data portability&quot;.&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Why is this important:&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Startups, but even unicorns tend to go bankrupt from time to time and the users don’t get to choose the way they say goodbye to their data, meaning that whatever you have aggregated during the months and years of using a service, &lt;a href=&quot;https://www.computerweekly.com/blog/Ahead-in-the-Clouds/What-the-Myspace-data-loss-debacle-tells-us-about-how-the-internet-values-creative-content&quot;&gt;might disappear as quickly as a MySpace backup&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The legal and technical challenge of data portability is widely unresolved, as the idea of the user leveraging their own data on different competing platforms does not seem to be a main concern for legislators who are definitely not influenced by &lt;a href=&quot;https://www.marketwatch.com/story/facebook-discloses-record-lobbying-spending-as-zuckerberg-braces-for-house-hearing-2019-10-22&quot;&gt;underpaid FAANG lobbyists&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The commoditization of the data ecosystem increases the possibility of creating data products, but it requires a governance for accessing accessing user data at least on a private and personal level.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;&quot; &gt;I&apos;m summarizing my experience as it might be useful to you also these days when some companies just fold like blam — think travel, mobility or any fluffed up unicorn. It’s about getting back what is legally yours, not necessarily about what you’re going to do with it.&lt;/p&gt;
&lt;h4 &gt;Big Tech:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.amazon.de/gp/privacycentral/dsar/preview.html&quot;&gt;Amazon&lt;/a&gt; a pretty new development, you have to confirm it in an e-mail, took 8 days. They provided 61 separate ZIP files one has to download one by one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://privacy.apple.com/&quot;&gt;Apple&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://takeout.google.com/?pli=1&quot;&gt;Google&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://account.microsoft.com/privacy&quot;&gt;Microsoft&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;Social:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.facebook.com/dyi/?referrer=yfi_settings&quot;&gt;Facebook&lt;/a&gt; &lt;code&gt;All of my data - JSON - High&lt;/code&gt; - &lt;code&gt;Create File&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.linkedin.com/psettings/member-data&quot;&gt;LinkedIn&lt;/a&gt; &lt;code&gt;Download larger data archive&lt;/code&gt; - &lt;code&gt;Request archive&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://twitter.com/settings/your_twitter_data&quot;&gt;Twitter&lt;/a&gt; &lt;code&gt;Twitter data download&lt;/code&gt; (No other download request is possible for 30 days. Like what.)&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;Media:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://support.deezer.com/hc/en-gb/requests/new?ticket_form_id=360000057869PR-#whatcanirequest&quot;&gt;Deezer&lt;/a&gt; One has to file a ticket.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Mixcloud&lt;/em&gt; You can&apos;t.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.netflix.com/account/downloadinfo&quot;&gt;Netflix&lt;/a&gt; You have to confirm in an e-mail.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://help.soundcloud.com/hc/en-us/articles/360004066174-General-Data-Protection-Regulation-GDPR-#whatcanirequest&quot;&gt;SoundCloud&lt;/a&gt; One has to file a ticket.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.spotify.com/uk/account/privacy/&quot;&gt;Spotify&lt;/a&gt; &lt;code&gt;Download your data&lt;/code&gt; You have to confirm in an e-mail.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;Location:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://foursquare.com/settings/privacy&quot;&gt;Foursquare&lt;/a&gt; &lt;code&gt;Export My Data&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.yelp.com/profile_privacy&quot;&gt;Yelp&lt;/a&gt; &lt;code&gt;Download a copy of your Yelp data&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 &gt;Code:&lt;/h4&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://github.com/settings/admin&quot;&gt;Github&lt;/a&gt; &lt;code&gt;Start export&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;If you don&apos;t find a product or service in the list below I kindly ask you to reach out to me and share your findings and experiences. Typically, the download mechanism can be found under &lt;code&gt;settings and privacy&lt;/code&gt;. In 100% of cases, it must be triggered manually. On average you will get feedback within some workdays. In 100% of the cases, it is only a manually downloadable, expiring compressed file.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>It's time to build data pipelines</title>
   <link href="https://dataengineering.academy/2020/05/04/its-time-to-build-data-pipelines.html"/>
   <updated>2020-05-04T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/05/04/its-time-to-build-data-pipelines.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Desperate times call for uplifting speeches from venture capitalists, apparently. Some six weeks into the COVID-19 induced economic crisis Marc Andreessen has blessed us with the glorious essay &quot;&lt;a href=&quot;https://a16z.com/2020/04/18/its-time-to-build/&quot;&gt;IT’S TIME TO BUILD&lt;/a&gt;&quot; (sic). &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;For those of you who have not read it yet, here&apos;s my brief rundown:&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The corona virus made us realise there are certain scalable products and services our society desperately needs but simply does not have enough of (surgical masks etc.), and the reason for that is not a lack of technical know-how or missing financial means for production, but that we consciously have chosen in the past not to build more of them. Andreessen is segueing from healthcare to housing, education and transportation to name some examples for sectors that we should literally expand both in scale and accessibility so we&apos;re better off collectively. After a spirited request to the political decision-makers on the right and the left he argues, although building is difficult, it will lead to the next wave of economic-idealogical prosperity in the US. Here&apos;s the finishing punchline:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Our nation and our civilization were built on production, on building. Our forefathers and foremothers built roads and trains, farms and factories, then the computer, the microchip, the smartphone, and uncounted thousands of other things that we now take for granted, that are all around us, that define our lives and provide for our well-being. There is only one way to honor their legacy and to create the future we want for our own children and grandchildren, and that’s to build.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset-h8g0zr.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;First and foremost, I feel the confusion of &lt;a href=&quot;https://vicki.substack.com/p/its-time-to-build-if-you-have-the&quot;&gt;Vicki Boykis, the queen of Normcore Tech&lt;/a&gt; (who also happens to be a master at curating artwork for her newsletter):&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;What are we, the people, supposed to make of an essay that is asking, nay, demanding that we build while at the same time ignoring what that the author has previously built - or, rather, invested in, has created as many, and maybe more, societal problems as it’s solved? The situation online as a result of many of these startups is so bad even &lt;/em&gt;&lt;a href=&quot;https://vicki.substack.com/p/data-centers-are-the-new-oil&quot;&gt;&lt;em&gt;the inventor of the internet&lt;/em&gt;&lt;/a&gt;&lt;em&gt; doesn’t like it anymore.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Let&apos;s put that very valid argument aside for a second and detach the message from the messenger. Furthermore, let&apos;s also ignore our current economic realities, the ideological landscape of the Western hemisphere and the systemic structure of our society. Let&apos;s ignore the cheap argument of going borderline Sheryl Sandberg (&lt;a href=&quot;https://www.cnbc.com/2019/01/20/sheryl-sandberg-admits-to-facebook-stumbles.html&quot;&gt;&quot;we need to do better&quot;&lt;/a&gt;) before the closing HATERS-WILL-SAY-IT&apos;S-PHOTOSHOP!-style pre-emptive defence at the end of the original post (&quot;I expect this essay to be the target of criticism.&quot;). &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Ben Thompson comes to &lt;a href=&quot;https://stratechery.com/2020/how-tech-can-build/&quot;&gt;a much more positive conclusion&lt;/a&gt; after putting the piece into a broader context:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;I do believe that It’s Time to Build&lt;/em&gt; &lt;em&gt;stands alone: the point is not the details, or the author, but the sentiment. The changes that are necessary in America must go beyond one venture capitalist, or even the entire tech industry. The idea that too much regulation has made tech the only place where innovation is possible is one that must be grappled with, and fixed. And yet, Andreessen himself said that we need to demand more from one another. We need to figure out how to fix Wisconsin, not flee from it. We need to figure out how to build real businesses that build real things, not virtualize everything. And we need to start fighting for not just infinite upside, but the sort of minute changes in cities, states, and nations that will make it possible to build the future. &lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Moving from individual action to cooperation to arrive at collaboration is a tremendous part of what made humankind prosper in the first place. Building is a manifestation of an agreement about what we as a species need for the future and how we should invest our collective resources and yes, it presumes a shared desire and a grand vision. Building as a process carries an inherent value that is able to give purpose and hopefully even meaning to our lives. In this context building becomes the enabler of prosperity.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/build.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;What does this have to do with data engineering, you ask?&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;A couple of weeks ago COVID-19 has derailed our timeline for starting our new school for data engineering. While we were trying to get used to the new rules of lockdown-life and attempting to reorganise our daily routines, we&apos;ve decided that there is no way that we can sit back and watch how this story unfolds without us doing anything about it. We are firm believers of the transformational power of education and that sharing our professional and personal experience in the realm of data is something that can help individuals and businesses improve their competences, which as a result &lt;a href=&quot;/2020/04/12/the-role-of-data-engineering-post-covid-19.html&quot;&gt;will make them more resistant to the economic impact of the crisis&lt;/a&gt; and possibly more attractive on the job market on the long run. So we came up with an idea that we&apos;re working on right now. I&apos;ll share more on this in a couple of weeks (sign up for &lt;a href=&quot;/&quot;&gt;the newsletter&lt;/a&gt; and follow us on &lt;a href=&quot;https://www.instagram.com/dataengineering.academy/&quot;&gt;instagram&lt;/a&gt; and &lt;a href=&quot;https://www.linkedin.com/school/pipeline-data-engineering-academy/&quot;&gt;LinkedIn&lt;/a&gt;).&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;And this is why the below paragraph of Marc Andreessen&apos;s essay has resonated with me so strongly, it comes from the same place as our intention. I am doing my best to avoid sounding overly pathetic and I won&apos;t compare our pseudo-altruistic relief effort to the ones essential workers are coping with on a daily basis. Still, I am excited about the next steps of this journey and encourage everyone to think about this for a minute:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;sqsrte-large&quot; &gt;&lt;em&gt;Every step of the way, to everyone around us, we should be asking the question, what are you building? What are you building directly, or helping other people to build, or teaching other people to build, or taking care of people who are building? If the work you’re doing isn’t either leading to something being built or taking care of people directly, we’ve failed you, and we need to get you into a position, an occupation, a career where you can contribute to building. &lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;h3 &gt;The &lt;em&gt;right&lt;/em&gt; time to build &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;It might sound controversial to some and a bit of a cliche to others, but I am actively trying to convince myself that &lt;a href=&quot;https://www.forbes.com/sites/johngreathouse/2020/03/11/coronavirus-recession-no-worries10-reasons-to-start-a-company-now/#4e3203902fa9&quot;&gt;starting a company during a recession is a good idea&lt;/a&gt;. There are &lt;a href=&quot;https://medium.com/swlh/13-massive-companies-that-started-during-a-recession-ba769e38d0ad&quot;&gt;plenty of well-known companies&lt;/a&gt; that made it and there are &lt;a href=&quot;https://www.profgalloway.com/wewtf&quot;&gt;honest founders&lt;/a&gt; to tell their heroic stories of success. On a separate but related note I&apos;d also like to introduce you to my old friend, &lt;a href=&quot;https://en.wikipedia.org/wiki/Survivorship_bias&quot;&gt;survivorship bias&lt;/a&gt;... but on the other hand:&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/93407227_2948055375279269_3312622779785805824_o.jpg&quot; alt=&quot;&quot;&gt;
            &lt;div class=&quot;image-caption&quot;&gt;&lt;p class=&quot;&quot; &gt;Source: &lt;a href=&quot;https://www.facebook.com/poppunkwayo/photos/a.141901315894703/2948055368612603/?type=3&quot;&gt;Facebook&lt;/a&gt;&lt;/p&gt;&lt;/div&gt;          &lt;/figcaption&gt;
        &lt;/figure&gt;    &lt;/div&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;&quot; &gt;In any case… back to building.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - April 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/05/03/the-data-janitor-letters-april-2020.html"/>
   <updated>2020-05-03T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/05/03/the-data-janitor-letters-april-2020.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://plausible.io/blog/remove-google-analytics&quot;&gt;Why you should stop using Google Analytics on your website&lt;/a&gt;&lt;br&gt;&lt;em&gt;Marko Saric, Plausible Analytics&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Good reasons, not even mentioning the real important ones.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://dfrieds.com/articles/data-science-reality-vs-expectations.html&quot;&gt;Data Science: Reality Doesn&apos;t Meet Expectations&lt;/a&gt;&lt;br&gt;&lt;em&gt;Dan Friedman&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Data &amp;amp; infrastructure have serious quality problems.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://blog.piekniewski.info/2020/04/13/deflaition/&quot;&gt;DeflAition&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/filippie509&quot;&gt;&lt;em&gt;Filip Piekniewski&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Scientist, Accel Robotics&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The realization that deep learning is not going to cut it with respect to self driving cars and many other applications is now an open secret.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.marksblogg.com/python-scraper-wireguard-vpn-ssh-proxy.html&quot;&gt;Python Web Scraping with Virtual Private Networks&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/marklit82&quot;&gt;&lt;em&gt;Mark Litwintschik&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, #BigData Consultant&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;I&apos;ll explore two solutions, the first using WireGuard and the second, using an OpenSSH SOCKS5 proxy.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@rbranson/10-things-i-hate-about-postgresql-20dbab8c2791&quot;&gt;10 Things I Hate About PostgreSQL&lt;/a&gt;&lt;br&gt;&lt;em&gt;Rick Branson, Engineering Leader, Segment&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In general I’d recommend starting with PostgreSQL and then trying to figure out why it won’t work for your use case.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@rakyll/things-i-wished-more-developers-knew-about-databases-2d0178464f78&quot;&gt;Things I Wished More Developers Knew About Databases&lt;/a&gt;&lt;br&gt;&lt;em&gt;Jaana Dogan, Engineer, Google&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;You are lucky if 99.999% of the time network is not a problem.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://insights.project-a.com/server-side-tracking-surprisingly-easy-83d1450cc08f&quot;&gt;Server-side tracking: surprisingly easy&lt;/a&gt;﻿&lt;br&gt;&lt;a href=&quot;https://www.linkedin.com/in/martin-loetzsch/&quot;&gt;&lt;em&gt;Martin Loetzsch&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Chief Data Officer, Project A Ventures&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Pixel-based tracking is dead.&lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://info.crunchydata.com/blog/optimize-postgresql-server-performance&quot;&gt;Optimize PostgreSQL Server Performance Through Configuration&lt;/a&gt;&lt;br&gt;&lt;em&gt;Tom Swartz, Software Engineer, Crunchy Data&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;If a query performs heavy joins or other expensive aggregate operations, or if a query is performing a full table scan where an index could be used, it will nearly always perform poorly, no matter how well the database settings are tuned.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>How to become a data engineer?</title>
   <link href="https://dataengineering.academy/2020/04/21/how-to-become-a-data-engineer.html"/>
   <updated>2020-04-21T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/04/21/how-to-become-a-data-engineer.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Ok, now that we&apos;ve spent some time with understanding &lt;a href=&quot;/2020/04/12/the-role-of-data-engineering-post-covid-19.html&quot;&gt;the current state of affairs when it comes to data engineering&lt;/a&gt; as a function or career choice, let&apos;s take a brief look at how to actually get there. &lt;/p&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;First off, not one of today&apos;s data engineers grew up as a kid imagining becoming a data professional once they grow up. Period. There is always a weird list of circumstances that led to a point when it became necessary to update their title on LinkedIn saying &quot;Data Engineer&quot; or something similar. For the last 10-15 years, the more prevalent the systematic extraction of value from data has become on a macroeconomic level, the more software engineers and data analysts have been moving/pushed towards setting up and maintaining relevant infrastructures that enable data products on the micro level. The undisputed (yet more and more often criticised) &lt;a href=&quot;https://hbr.org/2012/10/data-scientist-the-sexiest-job-of-the-21st-century&quot;&gt;hype around data science for the last 8 years&lt;/a&gt; has been accelerating the demand for a broader understanding of dealing with data within organisations and showed that a progressive approach is welcomed by shareholders as it indicates the intention and commitment for market dominance through innovation (just think how many times you&apos;ve heard about products with &quot;AI&quot; inside). When it comes to talent, tooling and organisational structure, the last years were all about the rapid professionalisation of the data ecosystem and as in any healthy market starting taking shape, many different competing offerings have emerged.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;For the last five years, there has been an ongoing discourse between industry stakeholders and the job market to try to codify the role and responsibility of a data engineer. With the expectations converging about what this position shall include on various levels, the educational market has started to take note and to try to address this need. However, just as quickly as the focus of data teams and the corresponding data tools change, the involved disciplines and their know-how have to adapt really fast as well. The adjustment of the supply side (available people on the job market with the right skillset) takes time though as it is influenced by so many other general economic and individual factors, and this seemingly endless yet rapid shapeshifting on both sides makes the supply-demand gap for data engineering as a function so difficult to define and grasp. As a benchmark, &lt;a href=&quot;https://pdfs.semanticscholar.org/6e7d/2195324c9293661fff547e8acdf80e31b7a6.pdf&quot;&gt;data science seems to be 5-6 years ahead&lt;/a&gt;. At &lt;a href=&quot;/&quot;&gt;Pipeline Data Engineering Academy&lt;/a&gt; we have a distinct point of view about the approach, general attitude and skillset a data engineer has to show in order to succeed, but for the sake of this post, allow me to address this in detail in a separate one.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Below you can find a non-comprehensive overview of the educational sector&apos;s sometimes outstanding and sometimes half-assed attempt of aiming at a moving target.  Be warned: you&apos;ll have to navigate between well-written, didactically thoughtful learning tools and more or less empty and pointless yet overpriced courses taking advantage of the overhyped expectations around a career in data. &lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset-1tmxc9.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;So you&apos;ve made the decision to become a data engineer &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;As mentioned above, there are various formal options promising to help you get there, so let&apos;s start with some self-assessment in order to understand which path would be the right one for you.&lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Motivation&lt;/strong&gt;: is data engineering a hobby you&apos;d like to explore, a pre-requisite for your sidegig, or the focus of your future career; is your motivation is intrinsic or extrinsic; would you like to improve your knowledge or to receive a certificate you can show off with etc.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Experience&lt;/strong&gt;: how new are you to the world of data and software&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Time&lt;/strong&gt;: think about the long-term commitment and the (quality) time available on a daily basis you can dedicate to learning&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Costs&lt;/strong&gt;: don&apos;t just consider the tuition or the subscription fees for courses, but the additional costs for housing, opportunity costs etc.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Learning methodology&lt;/strong&gt;: do you like to learn on you own pace by yourself, are you looking for a classroom experience with a lot of social interaction, would you like to focus on industry best practices or on a more academic approach etc.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; &gt;Keep in mind: there is no one-size-fits-all solution, but having an overview helps you to exclude the options that most likely won&apos;t work for you. If you have the chance to become an intern or an apprentice next to an experienced data engineer, you should probably take it - learning on the job offers the fastest learning curve from all the below options. Think strategically, invest your time and your money into the path that projects the highest ROI in your particular situation.&lt;/p&gt;
&lt;h3 &gt;Online courses&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Let&apos;s start with the online courses that offer an easy and affordable entry and self-paced learning paths for data engineers. Popular MOOCs like &lt;a href=&quot;https://careers.coursera.org/data-engineer/&quot;&gt;Coursera&lt;/a&gt; and &lt;a href=&quot;https://www.udemy.com/courses/search/?src=ukw&amp;amp;q=data+engineer&quot;&gt;Udemy&lt;/a&gt; offer learning tracks puzzled together mostly from an abundance of data science, data analysis, cloud engineering, Python and SQL programming courses, while &lt;a href=&quot;https://www.udacity.com/course/data-engineer-nanodegree--nd027&quot;&gt;Udacity is offering a nanodegree&lt;/a&gt; you can receive investing about 4-5 months of your time. This is a result of the above mentioned vague definition/understanding of what data engineering is and what it is supposed to be (at least that&apos;s what haters will say), and shows that thinking about data infrastructure and the non-fancy plumbing work one is confronted with later on is more like an afterthought rather than a real focus. With that being said, one has to show love for the quality of the courses, teachers and the materials: they can be engaging and rewarding, but by definition they will be far from what the industry considers state of the art. Easy to use websites, mobile apps with simple exercises (especially if you are looking into learning Python) all help to immerse yourself in the world of data related roles, but every time you hit the enroll button just ask yourself: are you in for the know-how or the certificate? &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.datacamp.com/tracks/data-engineer-with-python&quot;&gt;Datacamp&lt;/a&gt; and &lt;a href=&quot;https://www.dataquest.io/path/data-engineer/&quot;&gt;Dataquest&lt;/a&gt; are both specialised platforms focusing on the three career paths we consider traditional nowadays: data analyst, data scientist and data engineer. Heavily discounted subscription fees and extremely well made missions help you get started and build up the confidence to solve coding challenges you are faced with during a regular job interview process. Whether you prefer text-based or video-based learning, Python or R, whether you like to engage in community discussions or rather work by yourself should determine &lt;a href=&quot;https://www.coursereport.com/blog/dataquest-vs-datacamp&quot;&gt;which one&lt;/a&gt; &lt;a href=&quot;https://www.reddit.com/r/datascience/comments/85mhrc/datacamp_vs_dataquest/&quot;&gt;to pick&lt;/a&gt;. As a result of being much more domain-specific than generic MOOCs they offer more depth and cater for the future data professional significantly better (e.g. integration of Jupyter). For $200-$500 per annum you can get access to all their career paths so you can switch from engineering to analyst in case you deem the first one too challenging along the way.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;We can&apos;t ignore branded trainings organised by the key players in the data ecosystem: hardware manufacturers and SaaS providers offer extensive trainings so their clients can actually deal with their products (&lt;a href=&quot;https://cloud.google.com/certification/data-engineer&quot;&gt;Google Professional Data Engineer&lt;/a&gt;, &lt;a href=&quot;https://docs.microsoft.com/en-us/learn/certifications/roles/data-engineer&quot;&gt;Microsoft Azure Data Engineer&lt;/a&gt;, &lt;a href=&quot;https://education.oracle.com/certification&quot;&gt;Oracle Cloud&lt;/a&gt;, &lt;a href=&quot;https://web.archive.org/web/20200810001258/https://www.cloudera.com/about/training/certification/ccp-data-engineer.html&quot;&gt;Cloudera Certified Professional Data Engineer&lt;/a&gt; etc.). These courses are usually meant for engineers with a decent level of experience, and offer very specific know-how as opposed to a more general approach towards making things work. &lt;a href=&quot;https://www.pluralsight.com/&quot;&gt;Pluralsight&lt;/a&gt; offers a wide range of courses mainly focusing on B2B clients and their workforce, but offers a self-assessment tool to support you in picking the right courses matching your skill level and your goals.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;All fully online trainings lack the social component that defines traditional education: a multi-layered experience. The lack of direct in-person interaction with teachers, instructors, mentors and your peers turns the learning experience into a very efficient transfer of value, but it won&apos;t create a community in the traditional sense and it won&apos;t make you experience the social-emotional landscape of a school. Online trainings most likely won&apos;t enable you to become part of a (professional) network or an alumni group you can leverage later in life, and it&apos;s difficult for these platforms to provide proper career coaching as a result of the cultural and economic differences between regions their students come from (think job market in Silicon Valley vs. job market in Brazil).&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Recommended&lt;/strong&gt;: for newcomers trying to get a feel for what data professionals do and how disciplines compare, for software professionals without exposure to data products or statistics yet, BI managers and analysts broadening their competences and leveling up etc.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset-x2365p.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;Higher education&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Oh, the never-ending discussion about &lt;a href=&quot;https://www.forbes.com/sites/johnebersole/2013/12/12/an-upside-down-economy-education-cycle/#781ebe2147d4&quot;&gt;the countercyclical relationship between economic cycles and the offering and enrollment&lt;/a&gt; at higher ed institutions... Instead of going into detail about the reasons for disparities between geographic regions and socioeconomic structures in the sector of higher education, let&apos;s agree on a couple of fairly provocative statements: most universities are slow to adapt and tend to focus more on certification than the actual transfer of know-how. Historically, anticipating significant changes on the demand side of the job market and providing an appropriate response in form of an adjusted, market relevant curriculum has not been their strong suit. Hence we can observe that on the European continent - taking it as an example as it is the most relevant for Pipeline Academy - there are only a handful of universities offering Data Engineering in form of a course or a masters degree (MSc): &lt;a href=&quot;https://moseskonto.tu-berlin.de/moses/modultransfersystem/bolognamodule/beschreibung/anzeigen.html;jsessionid=44d6d308f00d712f7838582dfee2?number=61041&amp;amp;version=3&amp;amp;sprache=2&quot;&gt;Technische Universität Berlin&lt;/a&gt; (GER), &lt;a href=&quot;https://www.tum.de/en/studies/degree-programs/detail/data-engineering-and-analytics-master-of-science-msc/&quot;&gt;Technische Universität München&lt;/a&gt; (GER), &lt;a href=&quot;https://web.archive.org/web/20191020175312/https://hpi.de/en/projects/landingpages/master-of-science-in-data-engineering.html&quot;&gt;Hasso Plattner Institute&lt;/a&gt; (GER), &lt;a href=&quot;https://www.jacobs-university.de/study/graduate/programs/data-engineering&quot;&gt;Jacobs University&lt;/a&gt; (GER), &lt;a href=&quot;https://www.datasciencetech.institute/applied-msc-in-data-engineering/&quot;&gt;Data ScienceTech Institute&lt;/a&gt; (FRA), &lt;a href=&quot;https://web.archive.org/web/20191211215121/https://www.uni-potsdam.de/de/studium/studienangebot/masterstudium/master-a-z/data-engineering-master.html&quot;&gt;Universität Potsdam&lt;/a&gt; (GER) and &lt;a href=&quot;https://www.mff.cuni.cz/en/students/master-of-computer-science/4-degree-plans-software-and-data-engineering&quot;&gt;Charles University&lt;/a&gt; (CZ) to name the few who already integrated data engineering into their program.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;The pros and cons for enrolling at a university (time commitment, financial investment, degree as a signal for the job market etc.) is something I won&apos;t cover here, yet I encourage you to take a close look at the curriculum, the teachers and the alumni to get a sense of the quality of their education. Is an MOOC an alternative to the university experience? Well, it is not. But they can complement each other very well. Also, keep in mind that the unique landscape of colleges and universities (especially the top tier) in the US will sooner than later &lt;a href=&quot;https://www.profgalloway.com/post-corona-higher-ed&quot;&gt;become a playground of Silicon Valley&lt;/a&gt;, so expect a significant change in what we consider today a &quot;university experience&quot;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Recommended&lt;/strong&gt;: computer science and software engineer graduates with BSc., business and economics students looking for a specialisation that enables them to start their career at a tech company etc.&lt;/p&gt;
&lt;h3 &gt;Coding bootcamps and live trainings&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Software engineering is a trade without any strict professional standards (as opposed to medical doctors, architects, lawyers etc.). Up until now, the demand for software/data professionals on the job market has been constantly higher than the supply coming from formal education and this trend does not seem to change anytime soon. These two circumstances led to the emergence of coding bootcamps taking on the role of making certain popular professions in tech accessible for anybody (web developers, UX/UI designers etc.). First and foremost, they promise focus and speed for a fraction of the costs of a traditional education (depending on the duration, classroom size, reputation  etc. this will fall somewhere between $5.000-20.000 in the US, in Europe it’s around €6.000-16.000). They also offer a classroom experience with approachable instructors, a project portfolio that can be leveraged when looking for a new job, and of course career coaching that will make you nail job interviews.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Data science and it&apos;s little sister, the more approachable data analysis has been on the menu of bootcamps for a couple of years now, and it&apos;s popularity is indisputable. Data engineering however has been untouched for several reasons:&lt;/p&gt;
&lt;ol data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;The handful of experienced &lt;a href=&quot;https://web.archive.org/web/20191209165133/https://www.datacouncil.ai/blog/data-engineer-salaries-around-the-world-2019&quot;&gt;data engineers are currently hotter than ever on the job market scoring the highest salaries among data professionals&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Not all data engineers have the willingness and/or the capability to teach students.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Setting up a proper data engineering curriculum and figuring out the right admission process, approaches, tools, methodologies is an investment most bootcamps don&apos;t have the capacity to make.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset-yuvua2.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 style=&quot;text-align:center;white-space:pre-wrap;&quot;&gt;This is where Pipeline Data Engineering Academy comes in filling the gap. Our mission is to teach you how to build and maintain data products, machine learning systems and business intelligence tools through a 12-week intensive course. It is our belief that with this knowledge is going to set you up for a sustainable and rewarding long-term career.&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;But I digress... &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Before you sign up to our next training, I urge you to read &lt;a href=&quot;https://www.freecodecamp.org/news/coding-bootcamp-handbook/&quot;&gt;Quincy Larson&apos;s The Coding Bootcamp Handbook&lt;/a&gt;: it will help you prioritise your selection criteria and become more informed before you make a decision about joining any coding bootcamp. Approaching this journey with a certain level of scepticism is the healthy way forward. You probably want to take a look at student feedback on CourseReport and SwitchUp, the two leading coding bootcamp review sites: just make sure you consider &lt;a href=&quot;https://en.wikipedia.org/wiki/Survivorship_bias&quot;&gt;survivorship bias&lt;/a&gt; while reading through the comments. For the folks in the US, &lt;a href=&quot;https://www.thinkful.com/bootcamps/&quot;&gt;Thinkful has created a smart and simple bootcamp comparison tool&lt;/a&gt;. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;A note on COVID-19: most coding bootcamps have been forced to move their in-person courses online. There is a good reason why these institutions had classrooms in the first place, and this is something everyone expects to come back one way or the other after the corona restrictions are lifted. However, in case you choose to have a 100% remote coding bootcamp experience, you should ask yourself if the value for money ratio still holds up when compared to a recorded online training.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;strong&gt;Recommended&lt;/strong&gt;: for data analysts and jr. data scientists looking for expertise in building and maintaining data products, for engineers from different areas (frontend, devops etc.) transitioning into a new field etc.&lt;/p&gt;
&lt;h3 &gt;Other resources&lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;Whatever path you choose, there are a handful of resources that you should know about in order to complement your learning experience and to help you navigate your own ship. In case you find something you think should be on this list, please reach out to me and let me know!&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Notable articles and blog posts: &lt;/p&gt;
&lt;ul data-rte-list=&quot;default&quot;&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.dataquest.io/blog/what-is-a-data-engineer/&quot;&gt;What is a Data Engineer?&lt;/a&gt; (&lt;a href=&quot;http://dataquest.io&quot;&gt;dataquest.io&lt;/a&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://medium.com/@rchang/a-beginners-guide-to-data-engineering-part-i-4227c5c457d7&quot;&gt;A Beginner’s Guide to Data Engineering — Part I&lt;/a&gt;, &lt;a href=&quot;https://medium.com/@rchang/a-beginners-guide-to-data-engineering-part-ii-47c4e7cbda71&quot;&gt;Part II&lt;/a&gt;, &lt;a href=&quot;https://medium.com/@rchang/a-beginners-guide-to-data-engineering-the-series-finale-2cc92ff14b0&quot;&gt;Part III&lt;/a&gt; (&lt;a href=&quot;http://medium.com&quot;&gt;medium.com&lt;/a&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Data Engineering — Complete Reference Guide From A-Z [2019] (&lt;a href=&quot;http://towardsdatascience.com&quot;&gt;towardsdatascience.com&lt;/a&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.oreilly.com/radar/data-engineers-vs-data-scientists/&quot;&gt;Data engineers vs. data scientists&lt;/a&gt; (&lt;a href=&quot;http://oreilly.com&quot;&gt;oreilly.com&lt;/a&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Newsletter: Data Eng Weekly (&lt;a href=&quot;https://dataengweekly.com/&quot;&gt;https://dataengweekly.com/&lt;/a&gt;)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p class=&quot;&quot; &gt;Online mentoring platforms: &lt;a href=&quot;https://www.sharpestminds.com/&quot;&gt;SharpestMinds&lt;/a&gt;, &lt;a href=&quot;https://mentorcruise.com/&quot;&gt;MentorCruise&lt;/a&gt; &lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;&quot; data-rte-preserve-empty=&quot;true&quot; &gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;I encourage you to explore your options and start your transitioning into the world of data engineering. &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - March 2020</title>
   <link href="https://dataengineering.academy/the%20data%20janitor%20letters/2020/04/15/the-data-janitor-letters-2020-march.html"/>
   <updated>2020-04-15T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/the%20data%20janitor%20letters/2020/04/15/the-data-janitor-letters-2020-march.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;		 		 	 	 		&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://medium.com/starsky-robotics-blog/the-end-of-starsky-robotics-acb8a6a8a5f5&quot;&gt;The End of Starsky Robotics&lt;/a&gt;&lt;br&gt;&lt;em&gt;Stefan Seltz-Axmacher, CEO, Starsky Robotics &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The biggest open secret is that supervised machine learning does not live up to the hype. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://chartio.com/blog/why-we-made-sql-visual-and-how-we-finally-did-it/&quot;&gt;We Made SQL Visual - Why and How&lt;/a&gt; &lt;br&gt;&lt;em&gt;Dave Fowler, CEO, Chartio &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;We don’t pretend to have all the answers, and we’re not done innovating yet—eight in ten is not ten in ten—but it’s worth reflecting on what has worked so far. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20200423071251/https://blog.ptsecurity.com/2020/03/intelx86-root-of-trust-loss-of-trust.html&quot;&gt;Intel x86 Root of Trust: loss of trust&lt;/a&gt;&lt;br&gt;&lt;em&gt;Mark Ermolov, Positive Technologies &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Nobody should worry about security anymore if they are virtualizing on shared hardware. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.2ndquadrant.com/en/blog/postgresql-is-the-worlds-best-database/&quot;&gt;PostgreSQL is the worlds&apos; best database&lt;br&gt;&lt;/a&gt;&lt;em&gt;Kirk Roybal, 2ndQuadrant, PostgreSQL Trainer and DBA &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The title is not clickbait or hyperbole. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt; 	 &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The role of data engineering post-COVID-19</title>
   <link href="https://dataengineering.academy/2020/04/12/the-role-of-data-engineering-post-covid-19.html"/>
   <updated>2020-04-12T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/04/12/the-role-of-data-engineering-post-covid-19.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;The corona virus has changed consumer behaviour more significantly than any other event in this century and as a result, the job market has been turned upside down. The demand for certain products and services has vanished from one day to the other, but at the same time a handful of sectors are growing faster than ever before. While in the middle of setting up &lt;a href=&quot;/&quot;&gt;Pipeline Data Engineering Academy&lt;/a&gt;, we&apos;ve been caught up in this unexpected turmoil as well, so I thought to myself, the only right path forward is making sense of what&apos;s next to ensure that the value we aim to provide stays relevant.&lt;br&gt;&lt;/p&gt;&lt;h3 &gt;Tech is here to stay, but it&apos;s going to adapt &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;First, let&apos;s turn to &lt;a href=&quot;https://medium.com/humane-tech/12-things-everyone-should-understand-about-tech-d158f5a26411&quot;&gt;Anil Dash&lt;/a&gt; of Glitch for some clarity about what technology really is and how we would expect it to behave during a crisis:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;Technology isn’t an industry, it’s a method of transforming the culture and economics of existing systems and institutions.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;It is a misconception that tech is an industry by itself, therefore we can&apos;t just make predictions or assessments about the companies we consider part of the Silicon Valley ecosystem. All verticals are going to be impacted differently as consumption is quickly adjusting to the new rules of living under lockdown (i.e. travel vs. streaming media), while within those verticals winners will emerge on top of the losers corpses. Technology as a method, the main driver of growth in the last decade is forced to change, there is no way around it: if existing systems and institutions are adapting to the (post-)corona world, the approach of transforming their culture and economics is going to have to find it&apos;s new ways as well.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;&quot;Nobody ever got fired for buying IBM.&quot; Part #2&lt;/h3&gt;&lt;p class=&quot;&quot; &gt;During the time of economic growth and general positive attitude towards the near to mid term future, corporates and individuals are willing to take more risks when it comes to investments. The pre-corona world at the stock market made a large bet on growth driven by emerging technologies which require a significant financial investment and long-term commitment from organisations, but this type of capital also presumes consumer demand (in the form of disposable income of households) and a general positive trajectory of the markets. This notion resulted in a crazy number of startup unicorns and what felt like the next industrial revolution: things were good working in tech, job security was not a topic on peoples minds. &lt;a href=&quot;https://www.businessinsider.com/warning-signs-that-an-economic-bubble-is-about-to-burst-2019-5?r=DE&amp;amp;IR=T&quot;&gt;Some signals were there pointing towards an economic bubble waiting to burst&lt;/a&gt;, but being bitchslapped by a pandemic was not a scenario to be considered when discussing the upcoming fiscal year in the board room. &lt;/p&gt;
&lt;p class=&quot;&quot; &gt;In times of financial distress and unpredictability, people start looking at what they have and how this measures up when hardship is knocking on the door. Decision makers are going to have one thing on their minds: avoiding unnecessary risks. &lt;a href=&quot;https://www.theverge.com/2020/4/15/21222942/google-slowing-down-hiring-through-2020-covid-19-pandemic&quot;&gt;In Google-speak&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;&quot; &gt;&lt;em&gt;By dialing back our plans in other areas, we can ensure Google emerges from this year at a more appropriate size and scale than we would otherwise. That means we need to carefully prioritize hiring employees who will address our greatest user and business needs.&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;A significant part of the innovation economy is high-risk high-reward, and all of us are witnessing the process of sobering/falling out of love with unpredictable results - also referred to as &apos;market correction&apos;. Spending corporate money without demonstrating hard ROI will be nearly impossible, and this goes for hiring in particular.&lt;/p&gt;
&lt;blockquote&gt;&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://www.datanami.com/2020/04/03/how-covid-19-is-impacting-the-market-for-data-jobs/&quot;&gt;&lt;em&gt;Data scientists and data engineers have built-in job security relative to other positions as businesses transition their operations to rely more heavily on data... [ ]&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class=&quot;&quot; &gt;Makes sense, right? Basically, your tech job will be only as good as the value of your measurable output, and data related roles carry significant returns for consumers (external) and decision makers (internal) alike.&lt;/p&gt;
&lt;/div&gt;
                  &lt;img src=&quot;/images/posts/image-asset-siqilb.jpg&quot; alt=&quot;&quot;&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h3 &gt;The case for data and engineering&lt;/h3&gt;&lt;p class=&quot;&quot; &gt;One of the building blocks of technology itself that enabled/drove digital transformation with exponentially increasing pace is data, and &lt;a href=&quot;https://www.forbes.com/sites/forbestechcouncil/2018/05/02/building-a-data-driven-economy/#1c138432abde&quot;&gt;the players in various industries making more and more use of it&lt;/a&gt; (whether the economic value provided through leveraging technology and data is adequately measured by the stock market and company capitalisation shall remain a separate conversation). Taking a closer look at the various disciplines within the realms of a cross-functional digital product team (from agile coaches, design thinking facilitators, backend developers, UX designers, devops engineers etc.), it&apos;s likely that non-essential roles will have trouble flourishing as a consequence of the above when the main company concern is providing a cost-efficient and technologically robust core service, ideally based on data-driven decision-making. We are going to observe a general structural shift towards the maintenance and operations of existing tech products in exchange for developing new prototypes and &quot;failing quickly&quot;.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;It&apos;s back to the basics, and the mindset of an experienced data engineer is all about that. But don&apos;t take it from me, take it from &lt;a href=&quot;https://www.mckinsey.com/business-functions/mckinsey-digital/our-insights/how-chief-data-officers-can-navigate-the-covid-19-response-and-beyond&quot;&gt;McKinsey Digital recommending CDOs to focus on staying operational and ensuring business continuity&lt;/a&gt; as a first instance.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;Data engineering is one of the most underrated but essential roles in a digital product team that is going to stay on the hiring list of companies working with tech. Before COVID-19, we&apos;ve seen &lt;a href=&quot;https://techhub.dice.com/Dice-2020-Tech-Job-Report.html&quot;&gt;a 50% yoy growth in 2019 in the demand for engineers able to work with big data&lt;/a&gt; and a corresponding &lt;a href=&quot;https://hired.com/state-of-software-engineers&quot;&gt;rapid increase of salaries&lt;/a&gt;, and although the growth is somewhat slower now, the crisis verified that the demand is robust.&lt;/p&gt;
&lt;h3 &gt;This too shall pass. &lt;/h3&gt;
&lt;p class=&quot;&quot; &gt;If there is something the dotcom bubble and the 2008 economic crisis have taught us it’s that the economy is able to bounce back, it always has in the past. Whether it&apos;s a &apos;V&apos;-shaped return or more like an elongated &apos;U&apos;, we&apos;ll learn soon enough. Regardless, when it comes to the years between the high-times, you better stick to the basics... and it seems like data engineering remains one of the essential disciplines companies working with data desperately need.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt; &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - February 2020</title>
   <link href="https://dataengineering.academy/2020/03/08/the-data-janitor-letters-2020-february.html"/>
   <updated>2020-03-08T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/03/08/the-data-janitor-letters-2020-february.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt; 	 		 	 	 		&lt;/p&gt;
&lt;h4 &gt;
&lt;a href=&quot;https://doriantaylor.com/agile-as-trauma&quot;&gt;Agile as Trauma &lt;br&gt;&lt;/a&gt;&lt;em&gt;Dorian Taylor &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;The Agile Manifesto is an immune response on the part of programmers to bad management. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://scottlocklin.wordpress.com/2020/02/21/andreessen-horowitz-craps-on-ai-startups-from-a-great-height/&quot;&gt;Andreessen-Horowitz craps on “AI” startups from a great height&lt;br&gt;&lt;/a&gt;&lt;em&gt;Scott Locklin &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Machine learning will be most productive inside large organizations that have data and process inefficiencies. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://medium.com/@squarecog/five-interesting-data-engineering-projects-48ffb9c9c501&quot;&gt;Five Interesting Data Engineering Projects&lt;/a&gt; &lt;br&gt;&lt;em&gt;Dmitriy Ryaboy &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;DBT, Prefect, Dask, DVC, Great Expectations. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://debezium.io/blog/2020/02/10/event-sourcing-vs-cdc/&quot;&gt;Distributed Data for Microservices — Event Sourcing vs. Change Data Capture&lt;/a&gt;&lt;br&gt;&lt;em&gt;Eric Murphy, Debezium &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Great sales speech, good tidbits, make up your own mind. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://github.blog/2020-02-14-automating-mysql-schema-migrations-with-github-actions-and-more/&quot;&gt;Automating MySQL schema migrations with GitHub Actions and more&lt;/a&gt;&lt;strong&gt;&lt;br&gt;&lt;/strong&gt;&lt;em&gt;Shlomi Noach, Principal Software Engineer, GitHub &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Your mileage may vary, a reasonable case study. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://adguard.com/en/blog/ad-blocking-history.html&quot;&gt;Are ad blockers doomed or have we already won? A history lesson&lt;/a&gt;&lt;br&gt;&lt;em&gt;Vasily Bagirov, AdGuard &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Our guess is that to stay competitive, ad blockers will have to move to a network-level approach. It means that an ad blocker of the future will have to monitor traffic of the entire network.&lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 
 <entry>
   <title>The Data Janitor Letters - January 2020</title>
   <link href="https://dataengineering.academy/2020/02/16/the-data-janitor-letters-2020-january.html"/>
   <updated>2020-02-16T00:00:00+00:00</updated>
   <id>https://dataengineering.academy/2020/02/16/the-data-janitor-letters-2020-january.html</id>
   <content type="html">&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;p class=&quot;sqsrte-large&quot; &gt;Data engineering salon. News and interesting reads about the world of data.&lt;/p&gt;
&lt;p class=&quot;&quot; &gt; &lt;/p&gt;
&lt;blockquote&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;Code reading is the single largest expense in software development, and very few even talk about it. &lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://twitter.com/girba/status/1221519863904198658&quot;&gt;&lt;em&gt;@girba&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://web.archive.org/web/20200123161254/https://towardsdatascience.com/stop-hiring-data-scientists-30514028e202&quot;&gt;Stop Hiring Data Scientists.&lt;/a&gt;&lt;br&gt;&lt;em&gt;Luke Posey, CEO, Spawner &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Your ROI is suffering from an inability to hire properly. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://twitter.com/Stiivi/status/1214758735739932672&quot;&gt;Flow chart: Do we need to deploy a new technology/system? &lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/Stiivi&quot;&gt;&lt;em&gt;Stefan Urbanek&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, advisor, Data Engineering Academy.&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;You bet. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://thegradient.pub/gpt2-and-the-nature-of-intelligence/&quot;&gt;GPT-2 and the Nature of Intelligence &lt;/a&gt;&lt;br&gt;&lt;em&gt;Gary Marcus, CEO, Robust.AI &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;In essence, GPT-2 has been a monumental experiment in Locke&apos;s hypothesis, and so far it has failed. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;http://blog.lareviewofbooks.org/provocations/neophobic-conservative-ai-overlords-want-everything-stay/&quot;&gt;Our Neophobic, Conservative AI Overlords Want Everything to Stay the Same&lt;/a&gt;&lt;br&gt;&lt;em&gt;Cory Doctorow &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;Ultimately, machine learning is about finding things that are similar to things the machine learning system can already model. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;blockquote&gt;
&lt;p class=&quot;&quot; &gt;&lt;em&gt;Are you using #postgres via #docker for mac? Have you ever noticed EXPLAIN ANALYZE slowing down your queries by like 60x? The important takeaway is that our modern stacks are incredibly complex and fragile. &lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;&quot; &gt;&lt;a href=&quot;https://twitter.com/felixge/status/1221512507690496001&quot;&gt;&lt;em&gt;@felixge&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://www.altinity.com/blog/2020/1/1/clickhouse-cost-efficiency-in-action-analyzing-500-billion-rows-on-an-intel-nuc&quot;&gt;ClickHouse Cost-Efficiency in Action: Analyzing 500 Billion Rows on an Intel NUC&lt;/a&gt;&lt;br&gt;&lt;em&gt;Alexander Zaitsev, Co-founder, Altinity &lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;A single ClickHouse server can be used to collect and monitor temperature data from 1,000,000 homes, find temperature anomalies, provide data for real-time visualisation and much more. Since it is a single server, setting it up, loading 500B rows and running sample queries is very easy. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://mineracaodedados.wordpress.com/2020/01/12/ml-security-countermeasures/&quot;&gt;Security in Machine Learning Engineering: A white-box attack and simple countermeasures&lt;/a&gt;&lt;br&gt;&lt;a href=&quot;https://twitter.com/flavioclesio?lang=en&quot;&gt;&lt;em&gt;Flávio Clésio&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Senior Machine Learning Engineer, MyHammer AG&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;After running a simple script based in using Scikit-Learn, I noticed there’s some latent vulnerabilities not only in terms of objects but also in regarding to have a proper security mindset when we’re developing ML models. &lt;/p&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;div class=&quot;sqs-html-content&quot;&gt;  &lt;h4 &gt;
&lt;a href=&quot;https://tech.marksblogg.com/fast-ip-to-hostname-clickhouse-postgresql.html&quot;&gt;Fast IPv4 to Host Lookups&lt;/a&gt; &lt;br&gt;&lt;a href=&quot;https://twitter.com/marklit82&quot;&gt;&lt;em&gt;Mark Litwintschik&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, #BigData Consultant&lt;/em&gt;
&lt;/h4&gt;
&lt;p class=&quot;&quot; &gt;My interest here is in seeing the performance differences between using PostgreSQL with a B-Tree index versus ClickHouse and its MergeTree engine for this use case. The performance gap in the hourly lookup rate favouring ClickHouse is off by an order of magnitude. &lt;/p&gt;
&lt;/div&gt;
</content>
 </entry>
 

</feed>
