Methodology

The First and Second Models are two strictly independent datasets: they remain separate everywhere on the site and are never merged by default.

To make searching easier, the program may ignore certain marks when comparing text. The displayed Quranic text remains faithful to the source supplied by the project’s author, with its original letters and marks.

Repetition and lexical resemblance: what is the difference?

Al-Zarkashi, in his work al-Burhân, defines “lexical resemblance” (al-mutashâbih al-lafẕî) as recounting the same story in different forms and sequences, a device frequently found in Qur'anic narratives and announcements.

He thus refers to the same Qur'anic statement recurring in similar forms, where the resemblance lies in differences of addition, omission, substitution or word-order inversion — which requires the reader who memorises the Qur'an to double-check and pay increased attention; reading specialists call this type “al-mushkil” (the ambiguous).

  • The “repeated” (al-mukarrar): the same word recurring without any variation at several places in the Qur'an, such as the verse ⁧﴿فَبِأَيِّ آلاءِ رَبِّكُمَا تُكَذِّبَانِ﴾⁩ [Ar-Rahman, 13] — identical, literal repetition.
  • The case where the meaning is repeated with slight, very close variations in wording — this is precisely lexical resemblance (al-mutashâbih al-lafẕî).
  • The case where only the meaning is repeated without the words resembling each other — such as the repetition of the prophets' stories in very different styles and wordings — which falls outside the scope of lexical resemblance.

Methodological note on words of similar form but different meaning

This work is based on cataloguing the words of the Qur'an according to their written form and their locations, keeping every location where a word resembles another in its shape or pronunciation, even if its meaning varies according to context, derivation, morphological structure or its position in the verse.

The reader may thus find in this project words of similar form that do not carry the same meaning at all their locations. This is neither a cataloguing error, nor an unintended repetition, nor a confusion between words: it is an intentional, methodical choice, because a Qur'anic word is never understood in isolation from its context, its construction, or the verse in which it appears.

Graphic resemblance therefore does not always imply a resemblance in meaning, just as a difference in vowelling, derivation or context can shift a word from one meaning to another. All these locations have been kept so that the research material remains complete, and so that both the reader and the researcher can reflect, compare, and return to the verses in their original context.

Thus, “⁧أَلَّفَ⁩” (allafa, “to unite”) and “⁧أَلْفَ⁩” (alfa, “thousand”) are close in written form, but their meaning differs: ⁧﴿وَأَلَّفَ بَيۡنَ قُلُوبِهِمۡۚ لَوۡ أَنفَقۡتَ مَا فِي ٱلۡأَرۡضِ جَمِيعٗا مَّآ أَلَّفۡتَ بَيۡنَ قُلُوبِهِمۡ وَلَٰكِنَّ ٱللَّهَ أَلَّفَ بَيۡنَهُمۡۚ إِنَّهُۥ عَزِيزٌ حَكِيمٞ63﴾⁩ [Al-Anfal, 63] — here, “allafa” means uniting hearts, creating harmony and affection between them; the meaning of uniting hearts is explicit in the verse.

In another verse, however: ⁧﴿وَلَتَجِدَنَّهُمۡ أَحۡرَصَ ٱلنَّاسِ عَلَىٰ حَيَوٰةٖ وَمِنَ ٱلَّذِينَ أَشۡرَكُواْۚ يَوَدُّ أَحَدُهُمۡ لَوۡ يُعَمَّرُ أَلۡفَ سَنَةٖ وَمَا هُوَ بِمُزَحۡزِحِهِۦ مِنَ ٱلۡعَذَابِ أَن يُعَمَّرَۗ وَٱللَّهُ بَصِيرُۢ بِمَا يَعۡمَلُونَ96﴾⁩ [Al-Baqara, 96] — here, “alfa” is a number (a thousand years), with no connection to the meaning of uniting.

Likewise, “⁧سُنَّة⁩” (sunna, “way, immutable law”) and “⁧سِنَة⁩” (sina, “drowsiness”) resemble each other graphically, but their meanings are clearly different: ⁧﴿سُنَّةَ ٱللَّهِ فِي ٱلَّذِينَ خَلَوۡاْ مِن قَبۡلُۖ وَلَن تَجِدَ لِسُنَّةِ ٱللَّهِ تَبۡدِيلٗا 62﴾⁩ [Al-Ahzab, 62] — here, “sunnat Allah” denotes Allah's immutable way, the constant divine law in His creation.

Whereas: ⁧﴿ٱللَّهُ لَآ إِلَٰهَ إِلَّا هُوَ ٱلۡحَيُّ ٱلۡقَيُّومُۚ لَا تَأۡخُذُهُۥ سِنَةٞ وَلَا نَوۡمٞۚ⁩...⁧255﴾⁩ [Al-Baqara, 255] — here, “sina” denotes drowsiness, in a verse that denies any sleep for Allah.

Similarly, “⁧أُمَّة⁩” (umma) changes meaning according to context: ⁧﴿إِنَّ إِبۡرَٰهِيمَ كَانَ أُمَّةٗ قَانِتٗا لِّلَّهِ حَنِيفٗا وَلَمۡ يَكُ مِنَ ٱلۡمُشۡرِكِينَ 120﴾⁩ [An-Nahl, 120] — here, “umma” denotes a man embodying by himself the qualities of goodness, standing in the place of an entire community in piety and obedience.

Elsewhere, it denotes a group of people: ⁧﴿وَلَمَّا وَرَدَ مَآءَ مَدۡيَنَ وَجَدَ عَلَيۡهِ أُمَّةٗ مِّنَ ٱلنَّاسِ يَسۡقُونَ⁩...⁧23﴾⁩ [Al-Qasas, 23] — here, “umma” denotes a group of people.

The same word can also mean a period of time: ⁧﴿وَقَالَ ٱلَّذِي نَجَا مِنۡهُمَا وَٱدَّكَرَ بَعۡدَ أُمَّةٍ أَنَا۠ أُنَبِّئُكُم بِتَأۡوِيلِهِۦ فَأَرۡسِلُونِ45﴾⁩ [Yusuf, 45] — here, “umma” denotes a period of time. Likewise, “⁧آية⁩” (aya) can mean a sign, a clear proof, or a verse of the Qur'an; here too, only the context can settle the matter.

The same applies to “⁧الرُّوح⁩” (ar-rûh): depending on context, this word can refer to revelation, to the angel Gabriel, or to the vital spirit of a human being — none of these occurrences should be reduced to a single meaning without examining the context of the verse.

This is why these words of similar form have been kept, despite their different meanings, for the following reasons:

  • Removing certain locations would reduce the catalogued Qur'anic material.
  • The reader needs to see all similar locations in order to compare and reflect.
  • The precise meaning never appears in the isolated word but in the verse, its context and its construction.
  • The difference in meaning despite graphic resemblance opens a door to linguistic, rhetorical and contemplative reflection.
  • What does not reveal its wisdom to the researcher today may become clear tomorrow, as linguistic and Qur'anic research expands, without ever claiming to settle what is not established by evidence.

The presence of words that look similar but differ in meaning is therefore an integral part of this project's method, not a classification flaw. The reader should always go back to the full verse, the surah name and the verse number, before basing any linguistic or semantic judgement on a word in the table.

Note: the search only covers words or text matching exactly and in isolation (“isolated” meaning not attached to another word or even a single extra letter). For example, a search for “⁧يعلمون⁩” (they know) does not include “⁧فيعلمون⁩” (and they know). A second clarification: some forms, such as “⁧فيعلمون⁩” (and they know), are not repeated in the Quran. However, the absence of a word as a separate entry does not necessarily mean that it is not repeated: this program retains the longest repeated segment. For example, “⁧فسيعلمون⁩” (then they will know) occurs twice, followed by “⁧من⁩” (who) in both cases. The program therefore lists “⁧فسيعلمون من⁩” (then they will know who), rather than “⁧فسيعلمون⁩” alone.

Why the conjunction waw is separated for indexing and analysis

Separating the conjunction waw (“⁧و⁩”) from the word that follows it, in this work, is in no way an alteration of the Qur'anic script (rasm), nor a call to write the Qur'an other than according to the established script of the Mushaf or the rules of Arabic spelling. In Arabic, the conjunction waw is normally written attached to the following word — such as “⁧والله⁩”, “⁧ويهدي⁩”, “⁧ومن⁩” — because it is a letter that is graphically linked to the following word.

But this project has an analytical and indexing nature, not the nature of a Mushaf transcription. This is why the separation of the conjunction waw has only been applied within the database and the search tables, so that it appears as an autonomous linguistic element, one that can be counted, whose locations can be tracked and whose relationship with what follows it can be studied. The displayed Qur'anic text remains preserved in its original form; this separation remains a technical tool for research and analysis.

Hiding the conjunction waw inside the word it is attached to would prevent distinguishing between two locations that differ in structure and scope: a word appearing alone, and that same word preceded by the conjunction waw. For example, searching for the phrase “⁧يهدي من يشاء⁩” may give ten results, while restricting the search to the form “⁧و يهدي من يشاء⁩” reveals that only five of these locations are actually preceded by the conjunction waw. This result is not a pointless repetition, but distinct analytical information, useful for studying context, linking verses and understanding the transition structure between meanings.

Similarly, the name “⁧الله⁩” appears in 2,393 results, while the form “⁧و الله⁩” appears in only 240 results. The first figure indicates the word's overall presence across the whole text; the second indicates its presence at the precise locations where it is linked by a conjunctive function. Neither piece of information should be dismissed on the grounds that the other has more results: each answers a different research question and carries its own significance.

The frequency of the conjunction waw, which reaches 9,546 locations in these tables, is not a reason to neglect it, but rather proof of its importance in the structure of the text. The conjunction waw is indeed one of the tools most closely tied to the construction of the sentence and the verse; it links words, structures and meanings, and helps mark coordination, contextual order or transitions between the elements of discourse depending on context. Counting it and separating it analytically therefore opens the way to a precise study, inaccessible if it remained merged into the following word.

This method has also revealed particularities that ordinary search does not bring out: notably, Surah At-Takathur (102) is the only surah where the conjunction waw does not appear according to the criteria of this analytical index. This observation is not in itself a definitive rhetorical judgement, but it offers a clear starting point for the researcher wishing to return to the surah and study its structure, context and linking devices.

This principle is not limited to the conjunction waw alone: very frequent words and particles do not lose their value because of their frequency. The name “⁧الله⁩”, pronouns such as “⁧هم⁩” and “⁧إلى⁩”, or words such as “⁧الذين⁩”, “⁧في⁩” and “⁧أن⁩” recur very often; yet cataloguing them and linking them to their contexts reveals precise nuances between their uses. The large number of results therefore does not make the search pointless: on the contrary, it makes classification and analysis more necessary and more rigorous.

This is why this project has adopted the analytical separation of the conjunction waw: to allow the reader to search for the word alone, the word preceded by the conjunction waw, or the conjunction waw itself, and to compare results clearly and verifiably. The goal is not to artificially inflate the figures, but to allow a methodical reading of the text, distinguishing the continuous orthographic script from the grammatical and semantic elements useful to scientific research.

Practical tome/page correspondence rules

The volume/page correspondence for the First Model follows this rule: in volume 1, the PDF page number equals the table page number plus 24 (an unnumbered 24-page introduction); in volumes 2 to 13, the two page numbers match exactly. The Second Model's rules will be documented separately once the exports are received.

The Basmala in this project

All scholars recognise “بِسْمِ اللَّهِ الرَّحْمَنِ الرَّحِيمِ” as a verse of the Quran within Surah An-Naml (surah 27, verse 30). Its status at the beginning of surahs, however, is a classic point of scholarly difference.

For the school of Imam Al-Shafi‘i and a group of scholars, the Basmala is the first verse of Surah Al-Fatiha; this is why it is numbered 1 in copies of the Quran based on the Medina or Kufa count, such as the widespread Hafs reading. For the schools of Imam Malik and Imam Abu Hanifa, it is not part of Al-Fatiha: the first verse is then “الْحَمْدُ لِلَّهِ رَبِّ الْعَالَمِينَ”, and the last verse is divided in two to keep seven verses.

At the beginning of the other surahs (it opens 113 of the 114, all except Surah At-Tawba), the Shafi‘i and Hanbali schools regard it as a verse, while the Maliki and Hanafi schools see it as a blessed formula that separates the surahs without being part of them.

In this project, “بِسْمِ اللَّهِ الرَّحْمَنِ الرَّحِيمِ” is counted as verse 1 of Surah Al-Fatiha, following the numbering of the Hafs reading. This numbering does not take sides in the difference between the schools.

First Model — table organised according to the order of the Qur'an

In this model, surahs and verses are classified according to the known order of the Qur'an, starting with Surah Al-Fatiha, then Al-Baqara, then the other surahs to the end of the Mushaf. Under each verse are grouped the results or the duplicate groups related to it, then the table moves on to the next verse according to the order of the Mushaf.

It is therefore easy, in this model, to follow the results according to their actual Qur'anic location: surahs appear in ascending numerical order, and verses follow their order within each surah. The order of the rows here is a sequential Qur'anic order, whose purpose is to link the search results to their original locations in the Mushaf.

Go to the First Model →

Second Model — table organised according to search results and matches

In this model, surahs and verses are not classified sequentially, from the beginning to the end of the Qur'an. Each group of results begins with a verse containing the searched word or segment, and the other verses belonging to the same group of duplicates or matches are then grouped below it, regardless of their actual location in the Qur'an.

The reader can thus move from one surah to another, then back to a previous surah, following the locations of similar words and the matching results, rather than the order of the surahs in the Qur'an. The order of the rows here is an analytical search order, whose purpose is to bring together in a single set all the locations linked to the same word or segment, even if those locations are scattered across distant surahs and verses.

It is therefore normal and intentional, in this model, for a surah number to move from an earlier rank to a lower one, or for the same surah to appear at several scattered points in the table: this signals no classification anomaly.

Go to the Second Model →