How to Navigate Infolanka Lanka News Without Losing Your Mind
Infolanka Lanka News is one of those Sri Lankan news portals that sits somewhere between a traditional newspaper site and a social media feed. It covers local politics, crime, business, and entertainment. The site runs on WordPress, uses aggressive ad networks, and moves fast when something breaks. If you are just trying to read an article, it works fine. If you are pulling data from it regularly, you will run into rate limits, broken endpoints, and pages that redirect through at least three shorteners before showing you the actual content. The site URL is straightforward. You can browse categories from the top navigation or search directly. The homepage loads a grid of recent posts with thumbnail images, category labels, and publication timestamps. Each post link goes to a slug-based URL structure that looks roughly like /infolanka-lanka-news/infolanka-lanka-news-article-slug/. Archives exist by month and year. There is no obvious sitemap displayed, but the WordPress setup does generate a standard XML sitemap at /sitemap.xml if you want to scrape the full index. I spent about two weeks last year pulling article data from this site for a media monitoring project. The first thing you need to know is that the API is not public. There is no documented endpoint, no auth key, nothing. Everything is front-end rendered HTML that you have to parse yourself. That is fine until you realize how heavy the page loads are. A single homepage can be over 2 megabytes once you factor in all the ad scripts and lazy-loaded images. My initial script took nearly 45 seconds per crawl. After stripping the unnecessary request types and caching the heavy assets, I got it down to roughly eight seconds. Not great, but usable.
Here is the practical approach I ended up using. First, set your user agent to something that looks like a regular browser. The site does basic bot detection and will serve a different, lighter version of the page to anything it flags. Second, respect the request rate. Two requests per second is about the sweet spot. Go faster and you start seeing temporary blocks that last anywhere from a few minutes to a couple of hours depending on how aggressively you push. Third, use a headless browser instead of a simple HTTP request library. The site loads a lot of content dynamically through JavaScript, and a plain GET request will miss most of the actual articles. Playwright with Chromium worked well for my use case. Puppeteer is fine too, but I found Playwright slightly more stable when dealing with the redirect chains this site throws at you.
Common Pitfalls When Scraping This Site
The biggest issue I ran into was with article metadata. The category information, author name, and sometimes even the headline are loaded through separate API calls that do not always resolve in the same time frame. If you try to parse the HTML immediately after page load, you will get null values for several fields. The workaround is to wait for network idle state or explicitly poll for the metadata elements to appear. I added a ten-second timeout with a retry loop and that solved about 90 percent of the data quality problems. Another thing nobody tells you about this site is that the URL structure changes without warning. I had a pipeline that had been running smoothly for months and then suddenly started returning 404s on about 15 percent of the URLs. The site had migrated to a new permalink structure without changing their internal links. If you are building anything that depends on historical URLs, archive them with timestamps so you can correlate which version of the site was live when. Otherwise you will end up chasing ghosts trying to figure out why a perfectly valid link from last month no longer works today. There is also the ad clutter problem. Infolanka Lanka News puts ads inside article content, between paragraphs, and in the sidebar. When you are parsing the text body, you will pick up sponsored content that looks identical to real articles. I built a filter that checks for specific ad class names and iframe patterns, which removed most of the noise. You still need to manually review the output because some ads are embedded as plain text blocks and the filter catches only about 85 percent of them. The remaining 15 percent are usually obvious if you know what you are looking for, but they do slip through occasionally.
Get the Full Details
One more thing worth mentioning is the mobile version. The site serves a different layout to mobile user agents, and the mobile HTML is actually much cleaner and easier to parse. If you do not need the desktop-specific content like certain sidebar widgets or related post modules, switching your scraper to a mobile user agent can cut your parsing time roughly in half. I switched my production scraper to a mobile agent and the success rate went up significantly because the mobile pages are less likely to break during updates. The tradeoff is that you lose some content that only exists on the desktop version, so it depends on what you actually need.
What This Tool Is and Is Not Good For
For reading news or browsing current events, Infolanka Lanka News is adequate. The articles are mostly timely, the coverage includes stories that other local outlets sometimes miss, and the search function works well enough for basic queries. For anything beyond casual use, however, you should expect friction. The site is not built for developers or automated consumers. There are no open APIs, the documentation is nonexistent, and the infrastructure shows the wear of a site that has been running long enough to accumulate technical debt without a dedicated engineering team to address it. If you need reliable news data aggregation from Sri Lanka, I would recommend pairing this with at least one other source. Sirasa News and Ada Derana both have more stable technical setups and better documentation for programmatic access. Using multiple sources also protects you against any one site going down or changing its structure unexpectedly. Infolanka can fill gaps, especially on local government and regional stories, but it should not be your only feed. The site does offer a newsletter subscription and a mobile app, which are worth considering if you just want the content delivered without doing any of the technical work yourself. The newsletter arrives daily with a curated set of stories and the app pushes breaking news alerts. Neither is perfect timing-wise, but they are considerably less headache than scraping the site directly. I ended up using both the scraper and the app because the app caught stories that my scraper missed due to the redirect issues I mentioned earlier. Sometimes the manual check is the only way to get complete coverage.
One final note about the comments section. It is active, but it is also where a lot of the less reliable information circulates. If you are using this site for research or fact-checking, treat comments as informal opinion unless you can verify them elsewhere. I learned that the hard way when a comment thread on one article contained what looked like official government statement text that turned out to be fabricated. The article itself was fine, but the comment section had a whole sub-thread of people sharing unverified claims. Worth keeping in mind if you pull user-generated content alongside the articles.
