Browser tab memory and large page extractions

I’ve finally identified the problem with large page extractions and failures… The browser tab memory grows dramatically during the scrape and drops extractions to a crawl with multiple page timeout errors.

Recently on a long page extraction of 600 pages, after about 75 pages, Chrome starts slowing down (page reloads), then page timeouts. A look at Chrome’s Task Manager shows that for the tab being used memory usage gets out of control. In this case, over 8GB and growing.

THE SOLUTION (but requires monitoring): When you first notice Chrome (or Edge) starting to reload a page, open a new tab. Then as soon as the content highlights for extraction, close it. This extracts the last content, then has UWS load the next page (and subsequent pages) in the open tab. This flushes the tab memory and page extractions resume like you’ve just started. Unfortunately, this might need to be done multiple times for a large list.

If there was a way to clear the tab memory during the page extraction to get rid of the “previous/next” tabs accumulating during a page extraction, it would fix all the problems with memory/sluggishness and fail page extractions.

Still hoping for failed page extractions to be re-run at the end of an extraction. Seems an error list could be added at the end of an extraction to accomplish this.

Please authenticate to join the conversation.

Upvoters
Status

In Review

Board
💡

Feature Request

Date

5 days ago

Author

MrKhaki

Subscribe to post

Get notified by email when there are changes.