0% found this document useful (0 votes)
11 views2 pages

AWS Data Pipeline and Metadata Management

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views2 pages

AWS Data Pipeline and Metadata Management

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

●​ AWS pipeline(Batch)

●​ Pushdown predicate.
●​ If we don't have a crawler how can we get metadata?
●​ To connect to power BI from S3 we can go through Athena.
●​ For the step function we have to make sure that it runs a wait for
the crawler to run completely. What happens if the crawler fails ?
what happens if half the file loads. The md5 hash would still
come. Would the number of the rows or even a single datapoint
change the hash of the md file.

●​ Tejas grp.
●​ Mockaroo
●​ Control table for incremental loading
●​ Clustering in snowflake
●​ Multi-user concurrency in snowflake
●​ How does it affect the business ? which inventory items
need to be restocked
●​ If the business just wants to see 3 months of data how are
you going to peek inside the data. (bucketing is important
for this)
●​ Don’t give a restrictive dashboard.
●​ How to auto scale the system.
●​ Always consider the 4 Vs
●​ SCD
●​ Can we have 2 layers in medallion—- we can directly have
a source connection….. Or if there is a data breach…. Or
the customer is well known then we can skip the read
heavy layer.
●​ Have separate notebooks
3. Ravi
●​ How every technology provides fault tolerance.
●​ cost
4. Shiv
●​ What questions can be asked to a client in a real project.
●​ Scalable
●​ CI/CD
●​ In the entire architecture where can you have the keys

functionality and demonstration

You might also like