Snapshot của apify/got-scraping: 768★ · TypeScript. HTTP client made for scraping based on got.
Tóm tắt dựng từ metadata GitHub của chính dự án — chưa có bài review TopGit. Trang sẽ tự động cập nhật khi bài review đầy đủ được xuất bản.
VÌ SAO CHƯA CÓ REVIEW
TopGit viết bài đầy đủ cho repo có nhiều sao nhất và được yêu cầu nhiều nhất. Trang này là snapshot trong thời gian chờ — xem README gốc ở tab READ ME.
After many years of development, we decided to deprecate the got-scraping package.
The package will no longer receive updates or support.
For new projects, we recommend using impit. impit is a modern, powerful, and flexible HTTP client with fetch API based on Rust's reqwest library. It provides a similar feature set to got-scraping, including browser-like request headers, proxy support, and more.
Got Scraping
Got Scraping is a small but powerful got extension with the purpose of sending browser-like requests out of the box. This is very essential in the web scraping industry to blend in with the website traffic.
Installation
$ npm install got-scraping
The module is now ESM only
This means you have to import it by using an import expression, or the import() method. You can do so by either migrating your project to ESM, or importing got-scraping in an async context
A non-primitive unique object which describes the current session. By default, it's undefined, so new headers will be generated every time. Headers generated with the same sessionToken never change.
Under the hood
Thanks to the included header-generator package, you can choose various browsers from different operating systems and devices. It generates all the headers automatically so you can focus on the important stuff instead.
Yet another goal is to simplify the usage of proxies. Just pass the proxyUrl option and you are set. Got Scraping automatically detects the HTTP protocol that the proxy server supports. After the connection is established, it does another ALPN negotiation for the end server. Once that is complete, Got Scraping can proceed with HTTP requests.
Using the same HTTP version that browsers do is important as well. Most modern browsers use HTTP/2, so Got Scraping is making a use of it too. Fortunately, this is already supported by Got - it automatically handles ALPN protocol negotiation to select the best available protocol.
HTTP/1.1 headers are always automatically formatted in Pascal-Case. However, there is an exception: x- headers are not modified in any way.
By default, Got Scraping will use an insecure HTTP parser, which allows to access websites with non-spec-compliant web servers.
Last but not least, Got Scraping comes with updated TLS configuration. Some websites make a fingerprint of it and compare it with real browsers. While Node.js doesn't support OpenSSL 3 yet, the current configuration still should work flawlessly.
To get more detailed information about the implementation, please refer to the source code.
Tips
This package can only generate all the standard attributes. You might want to add the referer header if necessary. Please bear in mind that these headers are made for GET requests for HTML documents. If you want to make POST requests or GET requests for any other content type, you should alter these headers according to your needs. You can do so by passing a headers option or writing a custom Got handler.
This package should provide a solid start for your browser request emulation process. All websites are built differently, and some of them might require some additional special care.
For more advanced usage please refer to the Got documentation.
JSON mode
You can parse JSON with this package too, but please bear in mind that the request header generation is done specifically for HTML content type. You might want to alter the generated headers to match the browser ones.
This section covers possible errors that might happen due to different site implementations.
RequestError: Client network socket disconnected before secure TLS connection was established
The error above can be a result of the server not supporting the provided TLS setings. Try changing the ciphers parameter to either undefined or a custom value.
apify/got-scraping có 768 sao GitHub — tải lại trang để xem số mới nhất, hoặc xem trực tiếp github.com/apify/got-scraping. TopGit phản chiếu số sao của GitHub nhưng không cam kết đến từng phút.
apify/got-scraping có phải mã nguồn mở không?
TopGit chưa ghi nhận license cho apify/got-scraping. Phần lớn repo public trên GitHub là mã nguồn mở, nhưng điều khoản khác nhau từng repo — mở file LICENSE để xác nhận.
apify/got-scraping có website riêng không?
TopGit chưa ghi nhận URL trang chủ cho apify/got-scraping. Phần README ở tab phía trên thường có link demo, hoặc xem mô tả GitHub của repo.
apify/got-scraping là gì?
apify/got-scraping (apify/got-scraping) là dự án TypeScript trên GitHub. Theo mô tả gốc: HTTP client made for scraping based on got.
Đọc thêm về apify/got-scraping ở đâu?
Trang TopGit này là một snapshot — tab "Readme" hiển thị nguyên văn README của repo (đã bỏ link, giữ ảnh). Repo GitHub ở github.com/apify/got-scraping là nguồn chính thức.
Đọc đầy đủ README ở tab phía trên.
Muốn nghe thêm một ý kiến về got-scraping?
Hỏi một AI đọc được trang này — một cú bấm là có ngay nhận định về got-scraping.