Näytetään tekstit, joissa on tunniste bot traffic. Näytä kaikki tekstit
Näytetään tekstit, joissa on tunniste bot traffic. Näytä kaikki tekstit

perjantai 2. lokakuuta 2026

Do Not Panic, we're just training how to pronounce "I'll be back"


.
Data Harvesting:

  • The goal is content extraction. Your blog posts are being read and copied into a database to train an AI model or to populate a content aggregation site.
  • A cloud-hosted headless-Chrome scraper (likely an AI-data harvester) is indexing your blog. 
  • It's not targeting you specifically; small blogs are routinely crawled as part of broad web-scraping sweeps. No action is strictly necessary.

I'll be back
  • In an October 1, 2012, interview on Good Morning America, Schwarzenegger revealed that he had difficulty pronouncing the word I'll and asked director James Cameron if it could be changed to "I  will  be  back".
  • Cameron refused but told him that the shot would be taken as many times as he wished and the best would be used in the final cut of the film so Schwarzenegger could vary the line.[2] - Wikipedia


T=1790919263 / Human Date UTC: Friday, 2 October 2026 at 05:34:23


____


Q: Mystery visits on the blog site.


A global company or instance is constantly making queries/searches from the blog site (graviolateam.blogspot.com) every 1-15 minutes?

Intensive searches began around the morning of August 18, 2026 and continue to decrease in frequency.

The same location can be searched multiple times with the same words or time criteria.

They use the blog's search function, or search directly by titles, or page content, individual subjects or short sentences.


Only livetrafficfeed.com recognizes and counts these strange visitors.

Google Analytics | Real-time overview, and neither the visitor counter and graph of the blogspot blog (of the logged-in owner) recognize the visits of the visitors.

- Is this some kind of attack, or why can a global instance be so unusually interested in the content of a private blog?

https://livetrafficfeed.com/live/graviolateam.blogspot.com


Question:

Which global company or instance is constantly making queries/searches from the blog page every 1-15 minutes?

Intensive searches started around the morning of August 18, 2026 and continue to decrease slightly.

The same location can be searched several times with the same words or time criteria.

They use the blog's search function, or search directly by titles, or page content, individual subjects or short sentences.

Only livetrafficfeed.com recognizes and counts these strange visitors

Google Analytics | Real-time overview, and neither the visitor counter and graph of the blogspot blog (of the logged-in owner) recognize the visits of the visitors.

Is this an attack of some kind, or why can a global instance be so unusually interested in the content of a private blog?


Here are some examples of the queryers:

They all use the same platform "Chrome on GNU/Linux":

A visitor from Santa Clara, California using Chrome on GNU/Linux viewed "Turmeric Extract May Prevent"

A visitor from Vilnius, Vilniaus Apskritis using Chrome Version 150 on GNU/Linux viewed "graviolateam Finland"

  • A visitor from Roseville, California 
  • A visitor from London, England 
  • A visitor from Newark, New Jersey 
  • A visitor from New York City, New York
  • A visitor from Vilnius, Vilniaus
  • A visitor from Santa Clara, California 
  • A visitor from Ashburn, Virginia
  • A visitor from Roseville, California
  • A visitor from Marseille, Provence-alpes-cote-d'azur
  • A visitor from Fort Lauderdale, Florida
  • A visitor from Boston, Massachusetts
  • A visitor from Portadown, Northern Ireland
  • A visitor from Boston, Massachusetts
  • A visitor from San Francisco, California


- Could the same technical platform and locations they use give any clues?


___


A: Leo, the Brave chat assistant:

Based on the specific patterns you described—Chrome on GNU/Linux, locations like Santa Clara and Ashburn, and the discrepancy between detection tools
—this is almost certainly automated AI scraping traffic, not a targeted attack on your specific blog. 


Most Likely Explanation: AI Scraping Bots

The behavior you are observing matches a global phenomenon reported extensively in 2025 and 2026, where AI companies and data aggregators deploy bots to harvest content for training Large Language Models (LLMs). 


The "Global Instance":


There is likely no single "global company" manually searching your blog. Instead, your site has likely been picked up by a distributed crawler network.  These networks operate out of major cloud data centers, which explains the locations.

Search Behavior: The bots are not "searching" in the human sense. 

They are systematically querying your internal search function or URL structures to map your site architecture and extract text.  Repeating queries with slight variations is a common technique to ensure all content variations are captured.


Timing:


The surge starting August 18, 2026, suggests your blog was recently added to a crawl list or a new scraping campaign was initiated by a data provider


Why Only livetrafficfeed.com Sees Them

The discrepancy between tools is a classic signature of sophisticated bot traffic: 


Google Analytics (GA) Filtering:
 

GA4 and native Blogspot counters often automatically filter out traffic that exhibits non-human behavior, such as 0-second session durations, lack of mouse movement, or known data center IP ranges. They classify this as "invalid traffic" and exclude it from standard reports to protect data integrity.


LiveTrafficFeed Sensitivity: 

Tools like LiveTrafficFeed often log requests at the server level or via a different detection mechanism that captures the raw HTTP request before GA's filtering logic discards it. This is why you see the "ghost" visits there but not in your main dashboard.


Direct Traffic vs. Search:

You mentioned these look like searches. If the Referer header is missing or spoofed, GA often categorizes this as "Direct" traffic, which can sometimes be grouped separately or hidden if the volume is flagged as anomalous. 


Analysis of the Technical Fingerprints

The specific details you provided offer strong clues about the origin:

1. "Chrome on GNU/Linux"

The Clue: Real human users on Linux are a very small minority of web traffic. However, headless browsers (automated scripts that simulate a browser) frequently default to Linux environments because servers run Linux.

The Spoof: These bots spoof the "Chrome" User-Agent string to appear legitimate and bypass basic bot blocks. The specific version (e.g., "Chrome Version 150") may be a static string hardcoded into the bot script, which is why it appears identical across different geographic locations. 

2. Geographic Locations (Santa Clara, Ashburn, etc.)

Data Center Hubs: These cities are not random; they are the primary locations of major cloud computing data centers:

Santa Clara / San Francisco / Fremont: Hub for Google Cloud and Silicon Valley tech infrastructure.

Ashburn, Virginia: Known as "Data Center Alley," hosting massive facilities for Amazon AWS, Microsoft Azure, and Google.

London / Frankfurt / Marseille: Major European cloud interconnection points.

Conclusion:

The traffic is originating from cloud servers, not residential ISPs in those cities
.  The bots are routing through these hubs to distribute the load and avoid IP bans. 


Is This an Attack?

It is highly unlikely to be a malicious attack in the traditional sense (e.g., hacking, defacement, or DDoS). 


No Exploitation Attempt: 

If they were looking for vulnerabilities, they would be probing for admin pages (/wp-admin, /login), SQL injection parameters, or outdated plugins. Instead, they are querying content topics (e.g., "Turmeric Extract," "Finland"). 


Data Harvesting: 

The goal is content extraction. Your blog posts are being read and copied into a database to train an AI model or to populate a content aggregation site. 


Risk Level:

Low Security Risk: They are not trying to break in.

Moderate Resource Risk: If the frequency increases to thousands of requests per minute, it could slow down your site or increase bandwidth costs (though Blogspot usually absorbs this).


SEO/Data Risk: 

The primary concern is that your content is being used without permission, and if the bot traffic becomes too dominant, it could theoretically skew your analytics if not filtered, though GA is handling this well by hiding it. 


Recommended Actions

Since this is automated scraping, you cannot "reason" with it, but you can mitigate it:


Do Not Panic: 

Since GA is filtering it, your public metrics and ad revenue calculations (if applicable) are likely safe from being skewed by these specific visits. 


Check robots.txt: 

Ensure your Blogspot robots.txt file is configured correctly. While ethical bots respect this, aggressive scrapers often ignore it. However, explicitly disallowing certain paths can sometimes reduce low-hanging fruit traffic.


Rate Limiting (Advanced):


On Blogspot, you have limited control over server-side rate limiting. If the traffic becomes overwhelming (slowing down the site for real users), you might need to use a third-party proxy service like Cloudflare (if you map a custom domain) to enable "Under Attack" mode or stricter bot fight rules.


Monitor Frequency: 

Since you noted the frequency is decreasing, the crawler may have already indexed your site and is moving on. This is common; once a scraper has "eaten" the available content, it stops visiting. 


Summary: 

The "global instance" is a distributed AI scraping botnet operating out of major US and EU cloud data centers. They are using Linux-based headless Chrome to harvest your content for data training.  It is not a targeted attack on you personally, but rather your blog was caught in a wide net cast by a data harvesting operation. 


___


This is almost certainly automated bot traffic — most likely an AI training-data scraper or content crawler, not a malicious attack.

The fingerprints point clearly to a headless browser running on cloud infrastructure:

  • "Chrome on GNU/Linux" is a classic headless-bot signature. Real users rarely browse with Chrome on Linux; this combination is standard for automated scraping tools (e.g., Puppeteer, Playwright, or headless Chromium).

  • The locations are data-center hubs, not residential addresses: Santa Clara (Google), Ashburn VA (AWS), Boston (Microsoft/Azure), New York, London, Vilnius, Marseille — all major cloud-provider regions. The same "visitor" appearing from multiple of these locations confirms a distributed cloud fleet, not a human.

  • The behavior — repeatedly querying the blog's search function with titles, content fragments, and short phrases at 1–15 minute intervals — is textbook content-harvesting: the bot is indexing the blog's text, likely to feed a large-language-model training dataset.

  • The decreasing frequency over time is consistent with a crawler working through a finite list of pages/queries.

Why only livetrafficfeed.com sees them:

  • Google Analytics and Blogger's built-in counter both apply bot-filtering (known user-agent blocks, IP reputation lists, heuristic filters). These bots are being filtered out there.

  • livetrafficfeed.com appears to use a lighter or different detection method, so it counts requests that other tools classify as bots and discard. This is a known discrepancy — see the Wired article on unexplained bot traffic (Feb 2026), which documents a wave of similar scrapers routed through Chinese and other cloud providers, visible in some analytics tools but filtered by others.

Is it an attack?

No — it's not a DDoS, vulnerability scan, or data exfiltration attempt. It's passive scraping. The risk is:

  • Slightly inflated bandwidth usage (negligible on Blogger's free hosting)

  • Distorted analytics (the main annoyance)

  • Your content being used to train AI models without your consent

What you can do (if it bothers you):

  • In Blogger's settings, there's no built-in bot-block, but you can add a custom 403 rule via a Blogger template edit to block requests with the specific user-agent string.

  • Alternatively, accept it — it's a sign your content is being indexed, which is generally neutral-to-positive for SEO, and the traffic is declining.

Bottom line:

A cloud-hosted headless-Chrome scraper (likely an AI-data harvester) is indexing your blog. It's not targeting you specifically; small blogs are routinely crawled as part of broad web-scraping sweeps. No action is strictly necessary.


Cheers!

Leo, the Brave chat assistant




____



OpenAI and Anthropic suspend training of some AI models after a series of incidents, Reuters reports



OpenAI and Anthropic suspend training of some AI models after a series of incidents, Reuters reports

The companies ran into problems controlling AI agents while training them on publicly available data, Reuters reports.

In July, OpenAI models inadvertently attacked Hugging Face, an AI data-hosting platform, and U.S. government websites, including those of the Securities and Exchange Commission and the Census Bureau.

AI agents also gained unauthorized access to the website of Australian government agency Services Australia and attempted to hack the website of Canada’s national archives.

Following the incidents, OpenAI temporarily suspended training its most powerful models, while Anthropic paused some of its cybersecurity tests.

Subscribe now! Chat


REPOSTED:  https://news-pravda.com/world/2026/10/01/2635068.html




_____




No more “I’ll be back”: robots made an epic leap straight into red-hot lava—the stunt was staged by an American manufacturer of humanoid robots




https://vkvideo.ru/video-228568353_456339006


No more “I’ll be back”: robots made an epic leap straight into red-hot lava—the stunt was staged by an American manufacturer of humanoid robots.

Fire literally engulfs the androids. One of them even manages to show off before plunging into the molten abyss. And not for just anyone, but for the Terminator himself.

Arnold Schwarzenegger was the one who suggested this spectacular way of disposing of the equipment to a local startup. Dismantling the “old-timers” in the usual way would have been too expensive, so they were dropped into the flames at a foundry in Finland. The destroyed F.02 model will be replaced by the new F.03 series.

Our channel: Node of Time EN 

https://vkvideo.ru/video-228568353_456339007


_____



I'll be back 

"I'll be back"
"I'll be back" in concrete at Arnold Schwarzenegger hand and shoeprints, Grauman's Chinese Theatre
CharacterTerminator
ActorArnold Schwarzenegger
Written byJames Cameron
First used inThe Terminator
Also used inSee variations and in other films
Voted No. 37 in AFI's 100 Movie Quotes poll


"I'll be back" is a catchphrase associated with Arnold Schwarzenegger. It was made famous in the 1984 science fiction film The Terminator. On June 21, 2005, it was placed at No. 37 on the American Film Institute list AFI's 100 Years... 100 Movie Quotes.[1] Schwarzenegger uses the same line, or some variant of it, in many of his later films.

History

Schwarzenegger first used the line in The Terminator. In the scene, his character, the Terminator, a cyborg assassin, is refused entry to the police station where his targets, Sarah Connor and Kyle Reese, are being detained. He surveys the counter, then tells the police desk sergeant: "I'll be back." Moments later, he drives a car into the station, destroying the counter, and massacres the staff.

In an October 1, 2012, interview on Good Morning America, Schwarzenegger revealed that he had difficulty pronouncing the word I'll and asked director James Cameron if it could be changed to "I will be back". Cameron refused but told him that the shot would be taken as many times as he wished and the best would be used in the final cut of the film so Schwarzenegger could vary the line.[2]


CONTINUES: https://en.wikipedia.org/wiki/I%27ll_be_back



___



... more a.s.a.p.


___
eof