Joe Reis, co-author of Fundamentals of Data Engineering, joins Mohamed Elsherif and Mohamed Hassan to unpack what data engineers actually do: move data from source systems through ingestion, storage, transformation, and delivery for analytics and machine-learning use cases. Reis explains why durable engineering concepts matter more than rapidly changing tools, and why teams should begin with business problems rather than familiar technologies. The conversation covers startup data strategy, cloud-cost awareness, managed services versus self-hosting, and the organizational realities that shape effective data teams. Reis also explores real-time systems, local analytics, generative AI’s promise and accuracy limitations, and practical career advice for aspiring data engineers.
Episode notes
• Why Joe Reis became a “recovering data scientist” and moved toward data engineering
• The motivation for Fundamentals of Data Engineering and its focus on lasting concepts over vendor-specific tools
• The data engineering lifecycle: source systems, ingestion, storage, transformation, and serving downstream use cases
• Organizational challenges versus technical challenges in data engineering
• Choosing technology from the business problem, not the “curse of familiarity” with tools such as Spark or streaming platforms
• Startup data strategy: focus on core metrics, use third-party services where appropriate, and avoid premature complexity
• Cloud-cost literacy, budget alerts, and understanding how cloud architectures affect spending
• The boundaries and collaboration among data engineers, software engineers, analysts, and data scientists
• Real-time data, local analytics, generative AI, self-service BI, privacy, and the limits of model accuracy
• Building data teams around organizational communication patterns and Conway’s law
https://josephreis.com/