5/27/22, 5:18 PM Scrape Configuration Validator (SCV++) (SCV++.
WebHome) - XWiki
Amazon Confidential
Scrape Configuration Validator (SCV++)
Primary Contact vaidyb (user) How do I change this value?
Last modified 1 year ago by vishasri.
Table of Contents
Scope/Objective
The objective of SCV++ is to serve as a gatekeeper for validating scrape data before ingesting the same into UEC. The intent is to surface issues with the scrape
configuration (if any) and use the feedback provided by SCV++ to re-run scrape and improve the data quality. This is key to ensuring high data quality on the
data that is being ingested to UEC.
How SCV++ functions
SCV++ can be run on sample data/entire selection based on the competitor selection size
A comparison is made between the live scrape data vs the latest month scraped data available in UEC to validate if there is any change in value for the
major datapoints
Post validation, SCV++ flags an item as SUCCESS/WARNING/ERROR based on the below mentioned scenarios
"SUCCESS" - When it passes all configured validation checks(mentioned below)
"WARNING" - When validation checks fails for minor data points such as our_price and offering_availabiity, whose values are expected to change
frequently
"ERROR" - When validation checks fails for major data points such as brand and UPC, whose values are not expected to change
Output file with list of crawled page IDs along with the type of error/warning is provided for site training associates to validate along with a error sum‐
mary file
The data validation checks include:
Difference against existing value for the same datapoint of same item in UEC
Semantic data quality checks
Character count limit checks
Regex checks - check for presence/absence of some characters in the value
Null/Empty/missing value check
How does SCV++ run (Old Process)
SCV++ runs automatically over the Sporc Jobs triggered either via Scraper UI or Sporc Crawl Jobs.
The ST will login to a beta host and run a script to consolidate failed/warning batches by providing the SPORC job id and the date range.
The consolidated file will be transferred to the local machine, 3-5 samples will be picked for each attribute-error type combination
The samples are manually validated and categorized as actionable (scrape reconfiguration) and un-actionable (false positives)
Once scrape reconfiguration is completed for actionable errors, the updated scrape configuration will be pushed to prod post CR approval
The items are re-scraped and validated again? And scrape audit stage is unblocked moving all items to EPC for further downstream validation.
Overriding : In case, if no action or re-configuration is required for the competitor, then ST should need to open the Sporc Job link and needs to override all
the batching within the date range(range for which you have validated the data) in the Scrape Audit Tab.
Reprocessing: Reprocessing is required when the error or warnings are actionable and reconfiguration requires for the competitor. In such cases competitor
changed configuration needs to be pushed into prod before reprocessing of the data. Do keep in mind that, we can only re-process 20 batches at a time due to
system-limitation.
[Link] 1/2
5/27/22, 5:18 PM Scrape Configuration Validator (SCV++) (SCV++.WebHome) - XWiki
Note : In case if you have higher number of batches failed in the Scrape Audit stage, it is not recommended to re-process the batches instead you can run a
adhoc scrape job for the same using scraper UI.
What are the latest advancements in SCV++(Current Process- Revised Process)
In the current state SCV++ do not need to be validated manually. We have UI for validating the batches, where the samples get collated automatically for the
domain within certain period of time and no overriding or reprocessing is needed with the current flow.
Please visit the Link below to know about the revised process.
[Link]
Tags:
[Link] 2/2