Web traffic is made up of both web browsers, as well as Internet robots. These robots can also be called bots or crawlers.
Internet robot (bot)
An Internet robot is a program that crawls websites for a specific purpose. For web operators, some of these purposes are beneficial, whilst others are not. We can therefore divide bots into harmful ones and useful ones. The number of bot visitors to a site can easily exceed the number of visits done from a web browser.
Malicious bots use web content only for their own purposes or try to attack the web in some way. One possible type of attack is a DoS attack, which aims to overload the web enough for it to crash or be very slow and unusable. These attacks are not uncommon in e-shops.
Our e-shop platform automatically blocks known malicious bots and has built-in protection against DoS attacks. For this reason, you will not notice these attacks or thenegative effects of harmful bots.
Good (useful) bots
These bots most commonly belong to search engines and marketing platforms such as Google, Facebook, Seznam, Bing and others. They crawl your site to index the site itself for search purposes, or to update information from the site.
Pageviews
A page view is considered as one single access to the web. In general, it is one visit to the page, or any part of the page, which is loaded when the visitor interacts with the site. This measurement is more precise than Google Analytics displays, as they do not count all page views, but only synchronous page views with a return code of 200, and only for page views done from a web browser.
Higher traffic (number of accesses) generates higher requirements for the general upkeep of the e-shop. Every individual approach comes at a cost. With increasing traffic, the e-shop must serve more users (accessing the web via browsers and robots) at the same time. Costs of operations grow non-linearly as products and stock increase. The more products and stock contained in the e-shop, the more performance-intensive each approach is. This also works with mass operations.
The difference between a page view via a web browser and from a robot
Web browsers and robots both identify themselves when visiting a website. This identification is called User-Agent. It contains information about the name and type of a browser or a bot, its version, and sometimes also other information. According to this value, it is in most cases possible to distinguish whether a page is viewed by a bot or from a web browser.
However, if there is an attacker (crawler) who wants to harm you or get data from your e-shop, he can nowadays easily order a service (robot) that crawls the web for you and at the same time this robot lets you know that it is a browser (sends a browser User-Agent). Simply put, such a robot disguises itself as a web browser.
Statistics
Access data can be monitored by primary admins and super admins on the statistics page in our administration.
Our system divides page views into the following categories:
- Web Browser page-views
- Administrators page-views
- GoogleBot page-views
- FacebookBot page-views
- BingBot page-views
- AppleBot page-views
- Otherbots page-views
This detailed breakdown can be seen by clicking on a specific day in the page view graph. After clicking, the graph also shows the distinguishing of accesses according to return status codes.
The codes are as follows:
- 2xx - Successful page display.
- 3xx - Redirect to another URL (301, 302).
- 4xx - Accesses to a non-existent URL (404) or unauthorized access (403).
Product and supply count
This is the average number of products for the whole month. We record the value of the number of products and the number of stocks every day. We use this average for the price of expansion packages so that the packages are not changed just based on some fluctuation (for example, an import of a new supplier, or an error in the supplier's data).
Additional packages
The price for expansion packages is always based on the number of e-shop accesses for the given month, combined with the average number of products and stock in that month. For example, the price for December is based on data from November, and the price for January is based on data from December.
Statistics in the Admin part monitor the real traffic and the calculated number of expansion packages that cover this traffic. The number of packages that will be charged depends on the actual traffic on the site. One additional package covers an increase in traffic of 50,000 web-browser page-views and 250,000 robot page-views. As an example, if a site attendance is 175,000 browser page-views higher than included in the base package, it will be counted as 4 additional packages.
How to restrict bot access
- You can adjust behavior on the side of the requester or on the shop side in admin, section Web access settings.
Googlebot
Googlebot is the web crawler used by Google to discover and index web pages on the internet. It is a software program that follows links on websites to find new pages and updates existing ones. Googlebot's main function is to help Google build a comprehensive index of the web, which it uses to return relevant search results to users. Googlebot crawls web pages by requesting the page from the server and parsing its content, including text, images, and links. The data collected by Googlebot is used to determine the relevance and ranking of a webpage in search results. Googlebot is constantly crawling the web to ensure that Google's search results are up-to-date and accurate.
Googlebot's crawl rate
Googlebot's crawl rate is the speed at which it crawls and indexes web pages on the internet. The crawl rate can vary depending on a variety of factors, including the website's server speed, the website's popularity, the frequency of new content being added, and the overall quality of the website.
Google adjusts the crawl rate of Googlebot for each website individually, based on its perceived importance and popularity. For example, a high-traffic website with frequent updates will typically be crawled more frequently than a low-traffic website with static content.
You can adjust the browsing speed by setting the number of accesses per time period, called throttling. You can set the throttling settings (number of accesses per minute) directly in the administration in the Web Accesses section.
What are "non-credited" page views?
This number means the number of requests that are excluded from our pricing model.
It includes:
1. Requests made from IP addresses that are on our list of blocked addresses. Usually, these IP addresses are those from which made any attack in the past. One of the attack types is en.wikipedia.org/wiki/Denial_of_Service. These IP addresses are manually added by us. We exclude those page-views from the pricing so you don't pay for such requests.
We exclude those: a) at the top level of the infrastructure if the attacker attacks several shops (you cannot see such non-credited page-views in Statistics) or b) at the shop level if the attacker attacks only one shop (you can see such non-credited page-views in Statistics).
2. Redirects when action is made (e.g. adding product to a shopping cart) – HTTP code 302.
3. Forbidden requests or automatic DoS protection (HTTP codes 403 and 429, 444).
Non-credited HTTP response codes
244, 302, 403, 429, 444, 500, 501, 502, 503, 504, 505, 506, 507, 508, 510, 511
I have been notified by kvikbot 🤖 about high amount of pageviews on my online shop
First, you should read carefully and follow the instructions from kvikbot 🤖. If there is some unwanted traffic, you can investigate it further by downloading web access logs from the admin. It will give you a very detailed overview of the traffic.
For taking action regarding unwanted traffic, you can add several types of blocking rules in the admin interface "Settings -> Web Access" section.
If you are unsure about your website traffic or the next steps to take, feel free to reach out to our customer care at info@kvik.shop. We offer professional consultations on this topic.
I have been notified about increased robot page views.
These notifications are sent automatically if there is a spike in bot visits compared to the standard number of visits in previous days. In Settings -> "WAF Web Access," you can view bot visits for each day in a clear graph. If you click on specific days, you will also see the breakdown of visits by individual bots.
The basic web infrastructure package includes 1,500,000 bot visits per month (averaging approximately 50,000 bot visits per day). You can therefore use the graphs to determine whether it is necessary to limit traffic.
If you determine that you need to limit visits from a specific bot or bots, you can do so in the "WAF Web Access" section. At the bottom of the page (below the graph), you’ll find a list of individual bots and the option to limit their number of views, either for the entire month or per minute.