● AWS pipeline(Batch)
● Pushdown predicate.
● If we don't have a crawler how can we get metadata?
● To connect to power BI from S3 we can go through Athena.
● For the step function we have to make sure that it runs a wait for
the crawler to run completely. What happens if the crawler fails ?
what happens if half the file loads. The md5 hash would still
come. Would the number of the rows or even a single datapoint
change the hash of the md file.
● Tejas grp.
● Mockaroo
● Control table for incremental loading
● Clustering in snowflake
● Multi-user concurrency in snowflake
● How does it affect the business ? which inventory items
need to be restocked
● If the business just wants to see 3 months of data how are
you going to peek inside the data. (bucketing is important
for this)
● Don’t give a restrictive dashboard.
● How to auto scale the system.
● Always consider the 4 Vs
● SCD
● Can we have 2 layers in medallion—- we can directly have
a source connection….. Or if there is a data breach…. Or
the customer is well known then we can skip the read
heavy layer.
● Have separate notebooks
3. Ravi
● How every technology provides fault tolerance.
● cost
4. Shiv
● What questions can be asked to a client in a real project.
● Scalable
● CI/CD
● In the entire architecture where can you have the keys
functionality and demonstration