Challenges of data engineering in the AI era
As previously mentioned, data engineering is key to ensuring reliable data for AI, analytics and BI initiatives. Data engineers who build and maintain ETL pipelines
and the data infrastructure that underpins analytics and AI workloads face specific challenges in this fast-moving landscape.
Disparate data sources challenge most organizations: ISG predicts that by 2026, 8 in 10 enterprises will have their data spread across multiple cloud
providers and on-premises data centers that span multiple locations. This decentralization creates a dependency on specialized, siloed teams, inefficient pipelines,
development with high costs and slow time to value. As a result, data usage is limited and innovation is blocked.
Fragmented tooling: Most companies rely on multiple external tools and services or custom in-house solutions to manage their ETL processes. These
tools are often disjointed with different architectures, interfaces and outputs, which adds another layer of complexity to the data engineering experience and
leads to bottlenecked teams.
Handling real-time data: From mobile applications to sensor data on factory floors, more and more data is created and streamed in real time and
requires low-latency processing so it can be used in real-time decision-making.
Scaling data pipelines reliably: With data coming in large quantities and often in real time, scaling the compute infrastructure that runs data pipelines is
challenging, especially when trying to keep costs low and performance high. Running data pipelines reliably, monitoring them and troubleshooting when failures
occur are some of the most important responsibilities of data engineers.
Data quality: “Garbage in, garbage out.” High data quality is essential to training high-quality models and gaining actionable insights from data. Ensuring
data quality is a key challenge for data engineers.
Governance and security: Data governance is becoming a key challenge for organizations that find their data spread across multiple systems, with
increasingly larger numbers of internal teams looking to access and utilize it for different purposes. Securing and governing data is also an important regulatory
concern that many organizations face, especially in highly regulated industries.
These challenges stress the importance of choosing the right data platform for navigating constantly evolving data needs, particularly in the age of AI. But a data
platform in this age can also go beyond addressing just the challenges of building AI solutions. The right platform can improve the experience and productivity of
data practitioners, including data engineers, by infusing intelligence and using AI to assist with daily engineering tasks.
In other words, the new data platform is a data intelligence platform.
The Databricks mission is to democratize data and AI, allowing organizations
to use their unique data to build or fine-tune their own machine learning and
generative AI models to produce new insights that lead to business
innovation.
The Databricks Data Intelligence Platform is built on lakehouse architecture
to provide an open, unified foundation for all data and governance, and is
powered by a Data Intelligence Engine that understands the uniqueness of
your data. With these capabilities at its foundation, the Data Intelligence
Platform lets Databricks customers run a variety of workloads, from business
intelligence and data warehousing to AI and data science.
To get a better understanding of the Databricks Platform, the next page
shows an overview of the different parts of the architecture as it relates to
data engineering. The Databricks Data Intelligence Platform enables you to
execute all your data, analytics, BI and AI initiatives. As a 100% serverless
platform, it provides you with built-in features such as disaster recovery, cost
controls and enterprise security. Key components feature Mosaic AI with end-
to-end AI for both generative and classical AI; Databricks SQL, a serverless
intelligent data warehouse; unified data engineering with Lakeflow, a built-in
database with Lakebase; and AI/BI that integrates deeply with Databricks SQL
to easily extend business intelligence across your business.
6
UNIFIED AND OPEN GOVERNANCE WITH DATABRICKS
UNITY CATALOG
Unity Catalog (UC) provides a single, centralized data catalog for managing
permissions, auditing and lineage across all data and AI assets in the
lakehouse. It supports all major open formats, including Delta Lake, Apache
Iceberg™, Hudi and Parquet, so data engineers can choose the right format
for their workloads without lock-in. With built-in access controls, data quality
monitoring and column-level lineage, Unity Catalog makes it easier to enforce
compliance and trust while accelerating development. Interoperability across
formats and secure data sharing through Delta Sharing ensure engineers can
collaborate across teams, platforms and clouds while maintaining consistent
governance. Unity Catalog reduces complexity, scales with the needs of
modern data platforms and keeps openness and reliability at the foundation
of every data engineering workflow.
ACCELERATING PRODUCTIVITY WITH THE DATABRICKS
ASSISTANT
Databricks combines generative AI with the unification benefits of a
lakehouse to power a Data Intelligence Engine that understands your data’s
unique semantics. This allows the Databricks Platform to automatically
optimize performance and manage infrastructure in ways unique to your
business. Built on this engine is the Databricks Assistant — a context-aware
AI assistant that offers a conversational API to query data, generate code,
explain code queries and even fix issues. The Databricks Assistant helps
practitioners, including data engineers, become more productive by cutting
down the time it takes to build new data pipelines and to troubleshoot
pipeline issues.
BUILT-IN AND AI-READY DATABASE WITH LAKEBASE
Every application runs on a database and Postgres is the database of choice
for most application developers. Databricks Lakebase is a fully managed
Postgres database deeply integrated in the lakehouse to make it easy to build
data and AI applications. Lakebase automatically syncs lakehouse data and
features/models into an operational database for low-latency applications,
dashboards and CRM systems. With Lakebase, engineering teams can focus
on delivering insights to applications instead of building custom pipelines,
managing infrastructure or administering databases.
DATA INGESTION WITH LAKEFLOW CONNECT
Databricks enables organizations to efficiently ingest data from various
systems into a single, open and unified lakehouse architecture. Lakeflow
Connect offers simple data ingestion connectors for applications, databases,
cloud storage, message buses and more, into your lakehouse. These built-in
connectors provide efficient end-to-end incremental ingestion at scale, easy
setup with a point-and-click UI or API, and unified governance via Unity
Catalog. With Lakeflow Connect, Databricks has continued to innovate,
expanding the breadth of supported data sources with built-in, managed
connectors, as well as connectors with more customization options. Zerobus
Ingest enables you to push event data directly to your lakehouse without
requiring a message bus, thereby simplifying ingestion for IoT, clickstream,
telemetry and other similar use cases. Auto Loader is a fully customizable
connector that provides a Structured Streaming source for incrementally and
efficiently processing data as it arrives in cloud object storage, recommended
for usage with Declarative Pipelines.
DATA TRANSFORMATION FOR SPARK DECLARATIVE
PIPELINES
Spark Declarative Pipelines is a declarative ETL framework (built on the open
source Apache Spark™ Declarative Pipelines) that helps data teams simplify
and make ETL cost-effective in streaming and batch. Simply define the
transformations you want to perform on your data and let Spark Declarative
Pipelines automatically handle task orchestration, compute management,
monitoring, data quality and error management. Engineers can treat their
data as code and apply modern software engineering best practices like
testing, error handling, monitoring and documentation to deploy reliable
pipelines at scale. Declarative Pipelines fully supports both Python and SQL
and is tailored to work with both streaming and batch workloads.
PRODUCTION-GRADE DATA PIPELINES, NO-CODE
REQUIRED WITH LAKEFLOW DESIGNER
Analytics projects require collaboration between data engineers and business
analysts, who often work on separate platforms. While no-code tools can help
bridge this gap, challenges remain with siloed workflows, production
headaches and limited AI productivity. Lakeflow Designer solves this with a
no-code, AI-assisted tool — native to the Databricks Data Intelligence
Platform — that empowers business analysts and others to build production-
ready data pipelines with natural language. Powered by data intelligence,
Lakeflow Designer empowers shared, collaborative workflows with a built-in
path to production, built on AI that understands your business.
UNIFIED DATA ORCHESTRATION WITH LAKEFLOW JOBS
Lakeflow Jobs offers a simple, reliable orchestration solution for data,
analytics and AI on the Data Intelligence Platform. Lakeflow Jobs lets you
define multistep workflows to implement ETL pipelines, ML training workflows
and more. It offers enhanced control flow capabilities and supports different
task types and workflow triggering options. As the platform’s native
orchestrator, Lakeflow Jobs also provides advanced observability to monitor
and visualize workflow execution, along with alerting and troubleshooting
capabilities for when issues arise. Lakeflow Jobs comes with serverless
compute options, so you can leverage smart scaling and efficient task
execution.
PRODUCTION
Why data engineers choose Databricks
Lakeflow and the Data Intelligence Platform
So, how do the Data Intelligence Platform and Lakeflow help with each of the
data engineering challenges discussed earlier?
Unified data engineering experience: Lakeflow brings together all
key aspects of data engineering, from ingestion (Lakeflow Connect) to
transformation (Spark Declarative Pipelines) and orchestration (Lakeflow
Jobs), in a single solution built on top of the Data Intelligence Platform.
Through a centralized and simplified data platform, you can eliminate silos
across teams, minimize tool sprawl and improve operational efficiencies.
Efficient ingestion, wide range of data connectors: Lakeflow
allows you to efficiently ingest data, only bringing in new data or table
updates. With a growing set of native connectors for popular data sources, as
well as a broad network of data ingestion partners, you can easily move data
from siloed systems into your data platform. Ingesting and storing your data
in Delta Lake while leveraging the reliability and scalability of the Data
Intelligence Platform is the first step in extracting value from your data and
accelerating innovation.
Real-time data stream processing: Lakeflow simplifies
development and operations by automating the production aspects
associated with building and maintaining real-time data workloads. Spark
Declarative Pipelines provides a declarative way to define streaming ETL
pipelines, and Spark Structured Streaming helps build real-time applications
for real-time decision-making.
Reliable data pipelines at scale: Lakeflow uses smart autoscaling and
auto-optimized resource management to handle high-scale workloads. With
lakehouse architecture, the high scalability of data lakes are combined with
the high reliability of data warehouses, thanks to Delta Lake — the storage
format that sits at the foundation of the lakehouse.
■Data quality: High reliability — starting at the storage level with Lakehouse
Storage and coupled with data quality-specific features offered by Spark
Declarative Pipelines — ensures high data quality. These features include
setting data “expectations” to handle corrupt or missing data, as well as
automatic retries. In addition, both Lakeflow Jobs and Spark Declarative
Pipelines provide full observability to data engineers, making issue resolution
faster and easier.
■Unified governance with secure data sharing: Unity Catalog provides a
single governance model for the entire platform, so every dataset and
pipeline is governed consistently. Datasets are discoverable and can be
securely shared with internal or external teams using Delta Sharing. In
addition, because Unity Catalog is a cross-platform governance solution, it
provides valuable lineage information, so it’s easy to have a full
understanding of how each dataset and table is used downstream and where
it originates upstream.
Conclusion
As organizations strive to innovate by leveraging their data, data engineering
is a focal point for success by delivering reliable, real-time data pipelines that
make AI possible. With Databricks Lakeflow, built on lakehouse architecture
and powered by data intelligence, data engineers are set up for success in
dealing with the critical challenges posed in the modern data landscape. By
leaning on the advanced capabilities of Lakeflow and the entire Data
Intelligence Platform, data engineers don’t need to spend as much time
managing complex pipelines or dealing with reliability, scalability and data
quality issues. Instead, they can focus on innovation and bringing more value
to the organization.
In the next section, we describe best practices for data engineering and end-
to-end use cases drawn from real-world examples. From data ingestion and
real-time processing to orchestration and data federation, you’ll learn how to
apply proven patterns and make the best use of the different capabilities of
Lakeflow and the Data Intelligence Platform.