0% found this document useful (0 votes)
6 views19 pages

Data Observability Platform Proposal

The document outlines a proposal for a Data Observability Platform, detailing its features, advantages over Data Catalog Platforms, and its integration within a Post Modern Data Stack architecture. It emphasizes the need for a simpler, more agile solution tailored for smaller organizations, aiming to reduce costs and IT involvement while enhancing data engineering services. The platform includes key features such as data profiling, validation, lineage, expectation, and custom metrics, and is designed for rapid deployment and integration with existing systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views19 pages

Data Observability Platform Proposal

The document outlines a proposal for a Data Observability Platform, detailing its features, advantages over Data Catalog Platforms, and its integration within a Post Modern Data Stack architecture. It emphasizes the need for a simpler, more agile solution tailored for smaller organizations, aiming to reduce costs and IT involvement while enhancing data engineering services. The platform includes key features such as data profiling, validation, lineage, expectation, and custom metrics, and is designed for rapid deployment and integration with existing systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Data Observability Platform

(Specification, Design Details and Proposal)


What and Why?

Data Profile
Raw Data
Data Validation
Processed
Data Metadata
Data
Observability Data Lineage
Metadata Platform
Data Views

Data Quality Control


Metrics
Custom Metrics
How it can extend the data engineering (Advantage)?

Data Profile
Raw Data Advantage
Data Validation
Processed ● Data Expectation
Data ● Machine Learning
Metadata
Data ● Data Intelligence
Observability Data Lineage
● Data Integration
Metadata Platform ● Data Visualization
● Data Governance
Data Views ● Data Monitoring

Data Quality Control


Metrics
Custom Metrics
Data Observability vs Data Catalog Platform
● Data Observability Platform helps improving ROI
● Data Catalog Platform are part of COGS
● Data Observability Platform helps improve various data engineering services
● Data observability platform reduces various costs associated with data
engineering services i.e. intelligence, monitoring, and visualization
● Data Catalog Platforms are focussed mostly for data monitoring and data
governance and created mostly for metadata collection
● Data catalog is collection of data schema and its hierarchy
Modern Data Stack

Data Pipeline Data Storage Data Transformation Data Visualization Data Intelligence Machine Learning
Service Service Service Service Service & Data Science Service

Running in containers, each services has its own container(s), whole


MDS is either on cloud or on-prem
Modern Data Stack (MDS)

Data Pipeline Data Storage Data Transformation Data Visualization Data Intelligence Machine Learning
Service Service Service Service Service & Data Science Service

● Above services are running in cloud (SaaS) with SOA (Service Oriented Architecture)
● Above services are running in their own container(s)
● A full infrastructure management services needed to manage all the services
● The cost to setup a full scale MDS is very high due to infrastructure cost
● It will require large amount of resources to manage MDS
● The whole infrastructure take longer time to build and takes longer to return ROI
Post Modern Data Stack (Why?)
● Not every organization needs Modern Data Stack because organization size
and their data needs are very different
● Cost to setup MDS is very expensive and takes long time to set it up
● Organization takes very long time for ROI
● IT involvement is very high with the infrastructure setup and management

Conclusion:
Not every organization needs MDS, We need something very simple.
Post Modern Data Stack - Architecture

Data Pipeline Data Transformation API Service


Service Service

Data Storage
Service
Data Visualization Data Intelligence Machine Learning
Service Service & Data Science Service

Mostly all services are part of ONE single Service


Post Modern Data Stack (Specification)
● Smaller organizations needs are very different
● ROI is biggest reason to implement Post MDS architecture
● Reduce the role of IT to bare minimum or almost nothing
● Post MDS architecture are ready to run within a week or less
● All services are inside a single machine and only data storage is outside the
machine

Conclusion:
Post MDS should be small, agile and ready to go within a short amount of time.
Proposed Data Observability Platform (Key Features)

Data Profile Data Validation API Service

Data Storage

Data Lineage Custom Metrics Data Expectation

Mostly all services are part of ONE single Service


Data Observability Platform (5 main features)
● Data Profile
Read original
● Data Validation Only Store
source data ● Data Lineage data which
stored at S3 ● Data Expectation can create the
and RDS ● Custom Metrics listed features

Store structured
data into RDS
which is accessible
from API
Data Observability Platform integration with HSQ

S3 Data Observability Platform


Platform Features

RDS ● Data Profile RDS


HSQ ● Data Validation
Service ● Data Lineage
● Data Expectation
Customer Profile ● Custom Metrics
● API Service

API
Customer Profile
Future Extension
● Machine Learning
● Data Intelligence
● Data Integration
● Data Visualization
● Data Governance
● Data Monitoring
Data Observability Platform (Detailed Features)
● Data Profile
○ Data Overview
○ Data Interactions
○ Data Correlations
○ Outliers
● Data Validation
○ Data Errors and Warnings
○ Missing values in data
● Data Lineage
○ Row level lineage
○ Column Level Lineage
● Data Expectation
○ Set specific expectation inside the row and column based value(s)
● Custom Metrics
○ Add custom metrics based on data schema and complex SQL queries
Development Plan
● The current design supports 5 main features
○ Data Profile
○ Data Validation
○ Data Lineage
○ Data Expectation
○ Custom Metrics
● The architecture is developed on Post-Modern Data Stack
● An API service is added in the platform to provide external integration
● The platform can be developed independently HSQ development

If you like the plan we can discuss more about what is


added inside the actual implementation.
Resources (Other Data Observability Platforms)
● Big Eye
● Monte Carlo
● Accel Data
● [Link]
● Unravel Data
● [Link]
● Atlan
● Data Dog
● Datameer
● Soda Data
● Gigamon
Technology Expertise
I am expert in performing the following tasks individually or lead a team of engineers.

● Open-source Distributed ML platform, Data & ML Observability Platform


Development
● Architected and developed cloud-agnostic SaaS data analytics platforms with SOA
● Integrating various Machine learning and Artificial Intelligence engines and libraries
into modern and postmodern data stacks data platform
● Expertise to build fully functional prototypes using modern programming language
using Python, Go, Java, and React frameworks
● Cloud deployment Proficiency on AWS, Azure and GPC
● Setting up various platform building architecture including CI & CD tools,
● Services Oriented Architecture, MDM, Data Lake development
● Generative AI application development using large language models
● TechBio experience in Drug and Drug Interaction, AI Based Drug Discovery Libraries
and Platform
Tech Project Leader Experience and Expertise
I have completed the following tasks successfully as team lead or individual in various capacity.

● Developed Data and AI Business, AI Business Strategies, Delivered data


platform to F500
● Created and managed up to 30 members global engineering teams with
diverse backgrounds
● Managed Banking, Insurance, Credit Monitoring, & Healthcare customers up
to $2M ARR
● Delivered results to investors, CXOs, Cross-functional global teams, and
global customers
● Part-time AI, ML, Data and Startup Consultant for GLG, Guidepoint, AlphaSight,
& various venture capital companies
Project Availability
I am available to perform any one or the combination of the following tasks to assist any global organization or enterprise.

● Build end to end machine learning platform specific to customer requirements


● Applying various AI and ML libraries into existing enterprise data platforms
● Data Pipeline development to connect various data sources and store the
processed data to a cloud based storage platform i.e. RDS or Object Store
● Developing customer requirement specific prototypes
● Leading the development of data warehouse, data lakes and master data
management systems including KYC, custom data pipelines
● Security assessment for any existing enterprise platform deployment
● ROI improvement Guidance and implementation
● Data Platform Development, AI and ML training to corporate professionals
● Ad-hoc consultation
Avkash Chauhan
AI and TechBio Consultant
Email: avkash@[Link]
Phone: 650-713-9055
CONTACT [Link]

Common questions

Powered by AI

The Post-Modern Data Stack differs from the Modern Data Stack primarily in its simplicity and cost-effectiveness. The Modern Data Stack requires a substantial infrastructure investment with services running in cloud environments using SOA, resulting in high setup and management costs . In contrast, the Post-Modern Data Stack is designed for smaller organizations, requiring minimal IT involvement, with services consolidated into a single machine (except for data storage), which significantly reduces both the implementation time and costs . It is agile and can be deployed quickly, often within a week . This makes it more suitable for organizations with limited resources seeking quicker ROI return.

For small organizations, using a Data Observability Platform together with a Post-Modern Data Stack offers significant cost benefits by streamlining data infrastructure and reducing operational overhead. The Post-Modern Data Stack's design minimizes IT involvement and infrastructure complexity by consolidating services onto a single machine, which drastically lowers setup and ongoing maintenance costs . Combined with the Data Observability Platform, which reduces data-related inefficiencies through enhanced monitoring and validation, organizations can achieve increased ROI without the expenses associated with complex, large-scale data ecosystems . This combination provides small organizations with a scalable, efficient way to manage data processes while optimizing resource allocation.

The primary features of a Data Observability Platform include Data Profile, Data Validation, Data Lineage, Data Expectation, and Custom Metrics. Each feature contributes uniquely to data management: Data Profile helps in understanding data through overviews, interactions, correlations, and identifying outliers . Data Validation ensures data accuracy by detecting errors and missing values . Data Lineage tracks the origin and movement of data at the row and column levels . Data Expectation involves setting specific criteria for data values to meet certain conditions . Custom Metrics allow users to create tailored metrics based on data schemas and complex queries . Together, these features enhance data insights and governance.

Organizations transitioning from a Modern Data Stack to a Post-Modern Data Stack may face challenges such as managing the shift in infrastructure complexity and ensuring continuous data service without disruption. The Modern Data Stack involves a more extensive, distributed architecture with multiple services in separate containers, which may lead to initial resistance in adopting the more simplified Post-Modern approach due to potential perceived loss of functionality or control . Organizations must also carefully plan data migration and integration to ensure no loss of data quality and service capabilities during the transition. Additionally, addressing changes in IT roles and ensuring team buy-in can pose significant challenges, as the Post-Modern approach requires redefined responsibilities and new skill sets .

Data lineage in a Data Observability Platform enhances data reliability and governance by providing detailed tracking of data origins, movements, and transformations at both the row and column levels . This transparency helps organizations understand the data lifecycle and identify how data flows through systems, which is crucial for ensuring data accuracy and compliance with governance policies. Additionally, having detailed lineage information allows for better impact analysis and troubleshooting when data quality issues arise, thereby improving trust and integrity in data management processes .

A Data Observability Platform integrates with services like HSQ by allowing it to access and process data stored in systems such as S3 and RDS, facilitating comprehensive data profiling, validation, and lineage tracking . This integration provides benefits such as enhanced data quality controls and improved data workflow orchestration. By leveraging the data observability features, organizations can ensure the accuracy and completeness of data being manipulated and analyzed by HSQ, leading to improved data-driven insights and operational efficiencies. Additionally, this integration helps streamline data governance and compliance efforts by simplifying the monitoring and auditing of data processes .

Consolidating services into a single service in the Post-Modern Data Stack offers strategic advantages such as reduced complexity, lower setup and operational costs, and quicker deployment time. This architecture minimizes the need for extensive IT infrastructure management, as all services, except for data storage, are housed within a single machine . This reduces the role of IT in managing disparate systems, thus lowering maintenance and interoperability challenges. The simplicity and agility of this setup allow for faster adaptation to changing data needs and quicker realization of ROI, making it highly beneficial for organizations with limited resources and varying data demands .

In a Modern Data Stack, machine learning and data intelligence services play a crucial role in transforming raw data into actionable insights. These services can process large volumes of data to develop predictive models that aid in strategic planning and operational decision-making . By integrating machine learning into data pipelines, organizations can automate complex data processing tasks, improve the accuracy of data-driven insights, and drive innovation through advanced analytics . Data intelligence services further enhance value by providing tools for data visualization and interpretation, thus making insights more accessible to non-technical stakeholders. Together, these services support a comprehensive data strategy that leverages emerging technologies for enhanced business outcomes .

The use of Custom Metrics in a Data Observability Platform allows organizations to tailor metrics to their specific business needs, enhancing data-driven decision-making. By creating metrics based on unique data schemas and complex SQL queries, businesses can derive more relevant insights that directly address their operational and strategic objectives . This customization enables more accurate performance tracking against industry-specific benchmarks or internal goals, leading to informed decision-making processes. Moreover, Custom Metrics provide a scalable way to adjust analytics approaches as business parameters and data landscapes evolve, ensuring that decisions continue to be based on relevant and impactful data insights .

A Data Observability Platform enhances ROI for data engineering services by streamlining data management processes and reducing associated costs. It offers comprehensive observability features that enable proactive monitoring and validation of data pipelines, thus minimizing data errors and operational costs . It supports intelligent data integration, visualization, and governance, ultimately reducing time and resources spent on manual data oversight . Additionally, by focusing on quality control and maintaining data integrity, it helps in making better data-driven decisions, thus improving business outcomes .

You might also like