Dataset Viewer
Auto-converted to Parquet Duplicate
content_hash
string
document_id
string
title
string
text
string
url
string
kind
string
trust
string
quality
string
fetched_at
string
assessment_version
int32
topics
list
sources
list
000121753fabce9ecf08d738d7e15112d80582fad52799c519ca1fad710a1677
44b8b408a0fa0b78cdad746e
Publications | Directorate of Prisons
Directorate of Prisons Govt  of Tripura search Title | Prisons Occupancy 26-10-2025 | Prisons Occupancy 26-10-2025 Prisons Occupancy 27-10-2025 | Prisons Occupancy 27-10-2025 Prisons Occupancy 28-10-2025 | Prisons Occupancy 28-10-2025 Prisons Occupancy 29-10-2025 | Prisons Occupancy 29-10-2025 Prisons Occupancy 30-10-...
https://prisons.tripura.gov.in/publications?page=6
page
official_domain
full_text
2026-10-03T00:12:33.594398+00:00
3
[ "prisons" ]
[ { "document_id": "44b8b408a0fa0b78cdad746e", "url": "https://prisons.tripura.gov.in/publications?page=6", "title": "Publications | Directorate of Prisons", "kind": "page", "trust": "official_domain", "quality": "full_text", "fetched_at": "2026-10-03T00:12:33.594398+00:00", "assessmen...
00029842ca773dabb6422f2b3e8a12f081be47b674bb1bb4cc2a4bf541ab3e48
4f1b65ff57134e46e3a2b572
"https://sansad.in/getFile/lsscommittee/Joint%20Committee%20on%20the%20Waqf%20(Amendment)%20Bill,%20(...TRUNCATED)
"LOK SABHA\n\nREPORT OF THE JOINT\nCOMMITTEE ON WAQF\n(AMENDMENT) BILL, 2024\n\nEIGHTEENTH LOK SABHA(...TRUNCATED)
https://sansad.in/getFile/lsscommittee/Joint%20Committee%20on%20the%20Waqf%20(Amendment)%20Bill,%202024/18_Joint_Committee_on_the_Waqf_(Amendment)_Bill_2024_1.pdf?source=loksabhadocs
pdf
directory_verified
full_text
2026-09-30T08:22:59.148000+00:00
3
[ "government_statements", "legal_amendments_forms", "parliament", "religion_heritage" ]
[{"document_id":"4f1b65ff57134e46e3a2b572","url":"https://sansad.in/getFile/lsscommittee/Joint%20Com(...TRUNCATED)
0003c83b078684679c4142305a628e42b90fdb81d2d306cda600b14e051dc7f2
bd870ab4d36f1d77b54e8f17
"राज्य कृषि उत्पादन मण्डी परिषद उत्तर (...TRUNCATED)
"A+\nA\nA-\nहिन्दी\nमुख्य पृष्ठ\nहमारे बारे मे(...TRUNCATED)
https://upmandiparishad.upsdc.gov.in/Introduction.aspx
page
official_domain
full_text
2026-09-30T08:39:29.305769+00:00
3
[ "public_services" ]
[{"document_id":"bd870ab4d36f1d77b54e8f17","url":"https://upmandiparishad.upsdc.gov.in/Introduction.(...TRUNCATED)
0004b0112f4db1f8e1fe617c87c84903f05329851c96bd2f343035783ee8ae70
f39ea7b3d14b2c1c34fc63ba
Introduction | Manipur State Legal Services Authority (MASLSA) | India
"Share on Facebook\nShare of X (formerly Twitter)\nShare on Linkedin\nIntroduction\nManipur State Le(...TRUNCATED)
https://manipur.nalsa.gov.in/introduction/
page
official_domain
full_text
2026-09-30T05:38:22.530258+00:00
2
[ "legal_judiciary" ]
[{"document_id":"f39ea7b3d14b2c1c34fc63ba","url":"https://manipur.nalsa.gov.in/introduction/","title(...TRUNCATED)
0004d5fb41cb05b4a346ae15c09a113e04be92fea8252616377a6858cc6b3ebc
f28f185f2a7fd2034a2b59a8
"Appointment of Nodal Officers in Compliance with the Sexual Harassment of Women at Workplace (Preve(...TRUNCATED)
"Home\nDocument\nAppointment of Nodal Officers in Compliance with the Sexual Harassment of Women at (...TRUNCATED)
https://bishnupur.nic.in/document/appointment-of-nodal-officers-in-compliance-with-the-sexual-harassment-of-women-at-workplace-prevention-prohibition-and-redressal-act-2013/
page
official_domain
full_text
2026-09-30T21:35:41.029876+00:00
3
[ "legal_judiciary", "policy" ]
[{"document_id":"f28f185f2a7fd2034a2b59a8","url":"https://bishnupur.nic.in/document/appointment-of-n(...TRUNCATED)
0005569eaf3c54d3252f05b9b63e457fe52598ce49b002173b42a1f3e696dd8b
fa294a09a5d3078d93c2668b
Introduction | Kerala State Legal Services Authority | India
"Share on Facebook\nShare of X (formerly Twitter)\nShare on Linkedin\nIntroduction\nKeLSA is constit(...TRUNCATED)
https://kerala.nalsa.gov.in/introduction/
page
official_domain
full_text
2026-09-30T05:19:58.274491+00:00
2
[ "legal_judiciary" ]
[{"document_id":"fa294a09a5d3078d93c2668b","url":"https://kerala.nalsa.gov.in/introduction/","title"(...TRUNCATED)
00058e57613779119c9252e0668493745f13857d7c8802ed0768e4a275c9c010
8b892674d21fcd6cf27afb5e
"Accessibility Statement | Official Website of Department of Agriculture and Farmers Welfare, Govern(...TRUNCATED)
"search\nMain navigation\nMain Menu\nHome\nAbout Us\nAt a Glance\nAdministrative Structure\nAgricult(...TRUNCATED)
https://agri.tripura.gov.in/accessibility-statement
page
official_domain
full_text
2026-09-30T16:38:38.915227+00:00
3
[ "public_services" ]
[{"document_id":"8b892674d21fcd6cf27afb5e","url":"https://agri.tripura.gov.in/accessibility-statemen(...TRUNCATED)
0006c4f9ce4725e20e7207fb7b08bd0388f90b94478f5a967ba506a640c7f578
9eb58994243253c16df2f30f
Statistical Reports | District Court Solan | India
"Home\nDocuments\nStatistical Reports\nShare on Facebook\nShare of X (formerly Twitter)\nShare on Li(...TRUNCATED)
https://solan.dcourts.gov.in/document-category/statistical-reports/
page
official_domain
full_text
2026-09-30T04:32:48.694167+00:00
2
[ "legal_judiciary" ]
[{"document_id":"9eb58994243253c16df2f30f","url":"https://solan.dcourts.gov.in/document-category/sta(...TRUNCATED)
00071e0b8fbf8664be4a4103cf1e48384b49d0de16df22393e1f98b889c7ddd0
547723dc1b10630734029411
Former Judges | District and Sessions Court Chatra | India
"Home\nAbout Court\nFormer Judges\nShare on Facebook\nShare of X (formerly Twitter)\nShare on Linked(...TRUNCATED)
https://chatra.dcourts.gov.in/former-judges/
page
official_domain
full_text
2026-09-30T14:01:56.051245+00:00
3
[ "legal_judiciary" ]
[{"document_id":"547723dc1b10630734029411","url":"https://chatra.dcourts.gov.in/former-judges/","tit(...TRUNCATED)
00086b36831ba15689adb804db4a22a9a09c5f772cf449f435434224fa373815
0525380e7919225a6e53dc08
/docs_arch/11122VACANCY%20CIRCULAR06732620201028132005.pdf
"Page 1\nसत्यमव जबते\nकार्यालय प्रधान मुख्(...TRUNCATED)
https://www.incometaxhyderabad.gov.in/docs_arch/11122VACANCY%20CIRCULAR06732620201028132005.pdf
pdf
official_domain
ocr_text
2026-09-30T22:32:46.540966+00:00
3
[ "income_tax" ]
[{"document_id":"0525380e7919225a6e53dc08","url":"https://www.incometaxhyderabad.gov.in/docs_arch/11(...TRUNCATED)
End of preview. Expand in Data Studio

Bharat Guide: screened Indian public information documents

This snapshot contains 64,964 distinct normalized text bodies and 77,526 source records. Generated 2026-10-03T01:10:02.179230+00:00.

Contents and provenance

One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved as provenance, not counted as additional distinct documents. The canonical record is the lexicographically smallest eligible document identifier. Topic labels overlap.

Selection requires an eligible assessment whose body checksum still matches the current document, with full_text or ocr_text extraction. This is heuristic screening, not manual factual validation. Older and newer screening versions coexist and are reported in manifest.json. PTI secondary news, directory-only entries, summary-only records, rejected text and OCR awaiting review are not qualifying rows. A reviewed alias can be retained inside an eligible row's provenance, explicitly labelled. The text exports contain no raw response bytes or account credentials. The combined CSV labels screened and review documents explicitly; failure records are separate. The separate SQLite backup retains original evidence and operational records; its manifest records its own snapshot date.

Limitations

This is not a complete census of Indian websites, laws, schemes or facilities. Government-adjacent institutional publications are labelled separately from government sources. Retrieval time is not publication time. Source validity, amendments, legal force, policy deadlines and numerical OCR accuracy require checking the original. Exact normalized-body deduplication does not establish semantic equivalence.

Rights and use

Rights remain with each original publisher. Government provenance does not establish that every source is public domain or has one common reuse licence. No blanket licence is granted over collected source text. Consult each source's terms and applicable rights before redistribution or other reuse. Publication does not change the rights retained by each original source.

Releases

manifest.json gives checksums and row counts. publication.json tracks confirmed publication content hashes and the additional-document threshold. Each update is committed atomically with its data and manifest; the next release is due after at least 10,000 newly qualifying distinct bodies, not 10,000 fetched URLs or chunks.

Download layout and CSV fields

documents.csv.gz is sorted by the full url value in ascending, case-sensitive order.

documents.csv.gz is the main download: one deduplicated, nonempty extracted text body per row, combining screened and review content. The status column is good or review. Only good rows count toward the screened-document milestone. Review rows may include thin text, directory entries, news or incomplete extraction; keep them separate when evaluating quality. Duplicate source aliases are preserved in provenance/sources.csv.gz without repeating their text in the main CSV.

The CSV contains: content_hash, document_id, title, text, url, host, kind, trust, quality, state, category, fetched_at, metadata, sha256, status, assessment_version, assessment_status, assessment_reasons, topics. text is the full stored extracted body, not a summary. metadata preserves the representative document's SQLite JSON metadata. topics is a JSON list, and labels may overlap; unclassified documents are explicitly labelled. Fetch time is not publication time. A row represents a document body, not every relational SQLite row.

topics/document_topics.csv.gz maps topics to content hashes and good/review status; topics/summary.json provides counts. This index avoids copying full text into many folders. Topic labels are heuristic, not manual factual or legal validation.

fails/urls.csv.gz contains failure, robots-blocked, retry and OCR-pending records with their exact statuses; not every entry is a permanent failure. fails/empty_documents.csv.gz lists records without usable text. Neither is counted as a distinct text document. Unattempted queued URLs and seed inventories are not exported. No crawler handoff package is included.

CSV files use UTF-8, gzip compression and standard quoting for multiline text. Decompress to obtain an ordinary CSV, or read directly:

import pandas as pd
df = pd.read_csv("documents.csv.gz")
good = df[df["status"] == "good"]

Optional data/train-*.parquet files contain screened text in larger batches for the Hugging Face viewer. The CSV contains additional review content, so its row count is higher than the qualifying-document count. Source values are preserved; treat spreadsheet formulas in source text as untrusted, not executable instructions.

SQLite reference

SQLITE_SCHEMA.md lists every table and field; sqlite_schema.sql contains the reference schema. documents stores text and metadata; document_assessments stores screening decisions; document_topics stores labels. Original PDFs and responses are in raw_blobs.payload, linked through fetch_receipts.raw_sha256, with codec describing zlib compression or raw bytes. ocr_pages stores page-level results. These binary files and operational tables are not embedded in the CSV.

The private recovery files backups/bharat.sqlite and backups/sqlite-manifest.json retain their own snapshot date and checksum, which may precede the text release. The full SQLite backup includes crawl queues and source inventories even though this CSV release excludes them. Keep that backup private when publishing text-only data.

Sources and subject coverage

Sources include Indian central, state and local government portals; courts and legislative institutions; regulators and public-sector bodies; universities and research institutions; and other government-adjacent institutional sources. Domain endings alone do not establish ownership or reliability. Secondary news reporting, including PTI, is distinguished from official statements such as PIB releases.

This release has text provenance from 16,122 distinct website hostnames. Subdomains count separately. These are hosts represented in the released text, not a claim that every page of each website has been crawled. Failed-only, queued and seed-only hosts are excluded.

Website domain group Distinct hosts
.gov.in 11,598
.ac.in 1,791
Other .in 2,283
Other domain endings 450

The current classifier defines 36 broad subject labels, with the finer areas described below. The finer subject areas are descriptive; subtopic coverage is not counted independently. A website/document may appear under multiple topics, so topic totals must not be added together.

Topic host counts include all released source aliases for labelled text, including review material. Labels are automated and may be inherited across identical-text aliases. good means passed extraction/source/subject heuristics, not manual fact-checking. Examples name institution types and subject areas, not a verified inventory of every named institution.

Topic Website hosts Good documents Review documents Source types and finer subject areas
Law and judiciary 2,009 8,908 11,438 Courts, tribunals and legal-services bodies; Supreme Court, High Courts, judgments and legal aid
Schemes and benefits 3,471 10,755 4,590 Central/state scheme portals; benefits, eligibility, applications and beneficiary guidance
Policy and regulation 5,891 8,385 8,247 Ministries, departments and regulators; policies, regulations, guidelines and gazette notifications
Government statements 1,739 2,066 3,058 Official information offices and PIB; press releases, cabinet decisions and public statements
Universities and research 3,410 9,844 9,169 Universities, colleges and research institutions; admissions, curricula, accreditation and research
Religion, heritage and endowments 1,574 1,358 1,003 Heritage and endowment bodies; religious institutions, pilgrimage, archaeology and endowments
Mining and natural resources 2,412 3,367 3,623 Mining, geology and resource departments; minerals, coal, petroleum, water and forests
Tourism 1,491 3,567 3,417 Tourism departments and public tourism bodies; destinations, visitor information and travel advisories
Public finance and accounts 2,775 2,951 4,578 Finance departments, treasuries and audit bodies; expenditure, accounts, audits and appropriations
Population and statistics 1,045 972 1,008 Census and statistics agencies; population, demographics, household surveys and vital statistics
Banking and monetary policy 625 1,527 1,146 Banks and financial authorities; monetary policy, deposits, interest rates and financial stability
Other public services 3,742 8,463 7,448 Citizen-service and sector departments; agriculture, public health, sanitation, drinking water and transport
Constitution 209 149 125 Legislative and legal-information portals; constitutional text, provisions and constitutional materials
Acts, amendments, rules and legal forms 352 365 444 Legislative departments and legal repositories; Acts, amendments, repeals, rules and statutory forms
Hospitals and health facilities 188 995 1,435 Health departments and facility portals; hospitals, medical colleges, facility directories and infrastructure
Income tax 74 783 642 Tax authorities; direct taxes, returns, deductions and taxpayer guidance
Roads and public contracts 238 945 1,202 Roads departments and procurement portals; highways, tenders, contracts and awards
Public infrastructure 2,884 3,566 3,538 Public works and infrastructure agencies; railways, urban development and public projects
Digital infrastructure 395 1,401 744 Digital-government agencies; e-governance, broadband, BharatNet and digital public infrastructure
ONDC 4 13 20 Digital-commerce network information; ONDC and open-network participation
UPI and payment systems 18 5 13 Payment-system institutions; UPI, NPCI and payment-system documentation
Prisons and correctional administration 290 792 1,599 Prison and corrections departments; administration, jail manuals and aggregate reports
Juvenile justice 35 147 97 Child-welfare and justice bodies; legislation, observation homes and institutional reports
Reproductive health and rights 40 49 25 Health and rights-related public bodies; maternal health, family planning and pregnancy-termination frameworks
Minority rights 73 315 936 Minority-affairs departments and commissions; protections, programmes and public reports
Culture and arts 645 2,676 2,963 Culture departments and public cultural institutions; museums, arts and cultural programmes
Fire services and stations 227 374 321 Fire and emergency-service departments; stations, fire services and fire-safety guidance
Police services and stations 604 1,473 1,241 Police departments; stations, public services and departmental information
Credit and lending 77 167 175 Banks and credit-programme bodies; loans, guarantees, lending schemes and microfinance
Government grants and sponsorships 301 815 612 Grant-making and education departments; grants, scholarships and sponsorships
Government hospitality and hotels 59 647 720 Public tourism corporations and hospitality bodies; government hotels and state guest houses
Official contact directories 513 416 232 Official departmental directories; office contacts, telephone directories and who-is-who pages
International policy and statements 128 320 403 External-affairs and diplomatic institutions; bilateral relations, foreign policy and joint statements
Parliament and briefings 585 658 1,420 Parliamentary institutions; Lok Sabha, Rajya Sabha, briefings and parliamentary materials
Prime ministers: history and tenures 1 15 6 Official historical and institutional collections; former prime ministers and their tenures
Budgets 875 717 1,611 Finance and budget portals; budgets, demands for grants and appropriation documents
News reporting 1 0 11 Secondary news agencies, including PTI; news reporting kept distinct from official government statements
Unclassified 13,101 0 253,560 Extracted material without a subject label; retained for review, not assigned an invented topic

Agriculture currently sits within Other public services; it does not yet have a separately measured website count. Zero-count topics indicate tracked areas without labelled text in this snapshot, not completed coverage. Hospital records are not a reconciled national inventory; legal materials may include superseded versions.

For exact source URLs and titles, see provenance/sources.csv.gz. Document-level topics are in documents.csv.gz; the topic-to-document index and counts are in topics/. These files describe released material and do not include queued or seed-only website lists.

Downloads last month
458