1. About the System
System Development and Contributors
The system was developed by a linguistics research team led by Prof. Şükrü Halûk AKALIN and comprising Assoc. Prof. B. Tahir TAHİROĞLU and Lecturer Sinan YALÇINKAYA.
Artificial intelligence-supported development and analysis: The code development, optimization, designing definitions, meaning discovery is used in interpretation and documentation sessions. These are not the direct dictionary decision; the dictionary decision and final approval are controlled by an expert editor.
Technical infrastructure: The ClickHouse search table, Turkish BERT-based semantic vectorization, HDBSCAN clustering, local CPU worker pool, hot caches, Tailwind CSS and Chart.js.
The GTS Semantic Analysis System is a comprehensive, corpus-supported linguistic analysis platform designed for Turkish lexicographers and researchers. It brings together two primary data sources:
GTS — Advanced Turkish Dictionary
— entries, — meanings, — example sentences, and — proverbs. This is the core dictionary database developed as part of the Advanced Turkish Dictionary project.
T-BDLD — Turkish Context-Sensitive Lemmatization Corpus
A large-scale Turkish corpus containing approximately — tokens and — unique lemmas. Each word is lemmatized according to its context.
These figures can change in real time as the corpus is continually updated and expanded.
GTS combines dictionary data with authentic usage contexts to map living Turkish. By comparing these two sources, the system produces new-word candidates, new-meaning discoveries, derivational families, suffix-productivity measures, and contextualization patterns.
The Corpus Is Continuously Updated; Analysis Is Manual and Traceable
corpus (T-BDLD) It's constantly enriched and raisedThe new word candidates, meaning discoveries, and generational families, are updated and expanded in time.
Text corpus loading, cleaning and adding operations are monitored in the management interface; main production analysis can be started manually by the manager or will run automatically at the time set by Turkish time if the corpus update is open. In the current production order, the 03.30-hour mission operates the same profiling path as manual analysis, even if the master automation is turned off, even though it has its own key on; only if it is clearly asked to scan fully, the corpus will be reread. A live update bar at the top of the homepage And the analysis steps become visible:
- The corpus (T-BDLD) is read by local CPU worker pool...
- One-word candidates preparing...
- Calculating multiple-word candidates...
- Multiword surface forms are gathering...
- Calculating context index...
- Creating morphological families (This step takes especially long)
- Updating Binary cache and field/text distributions...
- Compiling results...
▪ At this stage the user Waiting for a while until analysis is complete When analysis is completed, all pages automatically start showing new results.
Not: Especially Processing morphological families It may take more than a few minutes, because there are steps to root out hundreds of thousands of formats, to build a chain of derivatives, additional productivity computation and falsification detection. The acceleration is done with the local CPU employee pool of the active server.
2. Glossary of Abbreviations
The acronyms and meanings that you will often see in the system:
Basic dictionary database in the system.
Modern Turkish use corpus, which is constantly growing, context sensitive, lematized.
Between 0 (100) and 100 (100), a point measuring the wealth of context in a word's collection.
It measures the tendency of two words to be seen together.
Normalized Point Controversial Information. PMI scaled between negative 1 and +1.
Positive PMI. Weak association from expected to zero; distinguishing word-type relationships.
It's a measure of frequency-resistant connectivity, theoretical upper bound 14.
A strong word in the relevant context, general corpus is a weighting measurer of whether it is distinctive.
The corpus passed is the word that can be considered, not on GTS, but in the dictionary.
Type of word: name, verb, adjective, envelope, context, pronoun, etc.
The language model (BERT) used for semantic similarities.
The density-based clustering algorithm used in the discovery of new meaning.
The parallel executive layer that shares heavy work with server processors, such as corpus reading, candidate collection, context and cache preparation.
The data is stored to open the most common analysis results quickly without recalculating.
Occurrence density in multi-word clusters of one word (0-100).
The addition to the new word derivative (-such as, c, - l, - you, -sal, -v, etc.).
The addition (-s, e, d, -den, -) that gave the word a grammatical role.
The roofs on the verb are: --strength, --ill, - --work.
3. Pages and How to Use Them
Each Page Has Its Own Guide
Top of all the pages listed below " What Is This Page? " This is the main directory panel. It can be opened by clicking, and you can get quick information about the purpose of that page, what it does, what special terms and useful clues are available. While this general manual covers the entire system, the page guides are specific to that page.
▸Home /
The entry page of the system. This is where GTS' main idea is summarized: The living map of Turkish is the dictionary, corpus and context in one place. The main sections, such as General Search, Morphology, corpus Statistics, Folders and Reports, are occurrence buttons; the bottom part is the occurrence buttons. 4 main list tabs and 17 solvent tools bulunur.
- Combined search proposals for GTS, corpus (T-BDLD) and new Word candidates
- Live list and discovery statistics of new Word candidates
- T-BDLD prevalence cards and cumulative graphics for proverbs, phrases and joint verbs
- Similar meaning search, analogy, semantic field, 3D map, network graph
- Meaning evolution, very meaningfulness, metaphor, entropy, context dynamics
▪ Homepage tabs and analyzer tools detailed description drink to this section Look.
▸Dictionary Search /search
From a single search box Six sources together search:
- Matching dictionary entry: The word found in GTS as dictionary headword
- Matching Meanings: GTS word in meaning/identification text
- T-BDLD Corpus: corpus Lemma or inflected form all word; no context is shown, only Lemma + inflelected form + lists often
- New Words: corpus removed candidate word
- New Meanings: Explored meaning extensions
- Reports: Keyword matching report pages
Not: T-BDLD Corpus Dictionary The management console is accessed from the session (verbologically defined dictionary entry); it is not included in the general search.
Turkish characters are auto-normalized.
If there are no results: If the wanted word is not defined in GTS in the form of a dictionary headword, the system searches for the closest match among the new word candidates and shows whether the candidate is waiting for the decision to be taken into the dictionary. The frequency cards in the dictionary enterry pages are supported by the prevalence of corpus (T-BDLD) and field/text distribution data.
corpus tab: The word you're looking for kalemler If it's a gravitational form like this, the lemma is connected (kalem) is listed below Attracted It's shown with his badge, whether the word is on GTS in every record. GTS or Yeni It's marked with a badge.
▸Computational Turkish Grammar /hesaplamali-turkce-dil-bilgisi
GTS Computational Turkish Grammar, ansiklopediden ayrı bir kamu yayın alanıdır. Türkçenin ses, biçim, söz dizimi, anlam ve kullanım düzenini bölüm bölüm; doğrulanmış dil bilgisi kaynakları, yeniden üretilebilir T-BDLD ölçümleri ve insan denetimli nötr tanıklarla açıklar.
- Türkçe, İngilizcenin veya başka bir dilin gramer kalıplarına uyarlanmaz; her kategori Türkçenin kendi ek sırası, biçim-ses ilişkisi ve cümle davranışı içinde çözülür.
- Ağır toplu ölçümler bölüm sürümünün derlem sıra numarasıyla sabitlenir; total sentence, sözcük ve tanık biçim sıklıkları ise sayfa açıldığında canlı derlem verisinden okunur.
- Okuma düzeni beş ana sekmedir: Temel çözümleme, Kuramsal modeller, Hesaplama ve ekler, İçgörü ve tanıklar ile Sorular ve başvuru. Canlı Türkçe Ek Atlası kendi içindeki Ek Envanteri, Bağlam, Morfotaktik, Kuramsal Model ve Canlı İçgörüler alt sekmelerini korur.
- Otomatik hesaplama yalnız örüntü keşfeder. Dil bilgisel genelleme, Türkçe uzmanlık kaynakları ve bağlam denetimi tamamlanmadan yayımlanmaz.
- Dört tamamlayıcı mercek birlikte, fakat birbirinin yerine geçirilmeden kullanılır. Graf kuramı lema-yüzey-ek-bağlam ilişkilerini; manifold yaklaşımı yüksek boyutlu biçim ailelerinin yerel komşuluklarını; kategori kuramı yüzeyden çözümlemeye ve lemaya giden dönüşümlerin bileşim tutarlılığını; yüksek dereceli mantık ise çözümleyici kurallar üzerinde geçerli Türkçeye özgü üst koşulları sınar.
- Yöntem karmasının yayın kapısı yalnız matematiksel uygunluk değildir. Graf ve manifold keşif adayı üretir; kategori ve mantık katmanları veri soyunu ve kural sınırlarını denetler. Gerçek derlem tanığı, kaynak yayılımı ve insan incelemesi bulunmadan sonuç Türkçe için yayımlanmış dil bilgisi kuralı sayılmaz.
- On iki hesaplama katmanı ayrı görev ve durumlarla raporlanır. Shannon entropisi ile MDL dağılım ve betimleme ekonomisini; Markov zinciri türlenmiş ek geçişlerini; cebirsel yarı halka ağırlıklı çözümleme yollarını; paradigma matrisi gözlenen ve beklenen hücreleri; karşı örnekler kural değişmezlerini; eşyazımlılık katmanı çok lemaya açılan yüzey ve bağlam çözümleme adaylarını; etkin öğrenme inceleme önceliğini; zaman dizileri gerçek belge tarihli üretkenliği; en uygun taşıma alanlar arası dağılım farkını; Transformer cümle içindeki bağlamsal ayrıştırmayı; Ek Üretkenlik İndeksi güncel derlemdeki türetim gücünü; Canlı Türkçe Ek Atlası ise yapım, çatı ve exponentialsnin sıklık, yaygınlık, bağlam ve morfotaktik birleşimlerini inceler.
- Ölçüm durumu kamu kartında açıkça yazılır. Sürümlü ölçüm gerçekten hesaplanmış sayıyı, etkin denetim yayın doğrulamasında çalışan kuralı, pilot ise veri ve yöntem sözleşmesi hazır olmakla birlikte henüz dil bilgisel bulgu olarak yayımlanmayan hattı gösterir.
- Entropi doğruluk puanı değildir. Bölümdeki Shannon entropisi aynı yöntemle çıkarılan yedi ek kümesinin ağırlık dağılımını ölçer. Koşullu entropi Markov geçişlerindeki kalan belirsizliği, entropi hızı ise uzun dizideki ortalama yeni bilgiyi inceler. Yüksek değer doğrudan üretkenlik, düşük değer de yanlışlık anlamına gelmez.
- Markov ve cebirsel çözümleme birbirini tamamlar. Markov modeli bir yolun tanıklanmış dizilere göre şaşırtıcılığını, ağırlıklı sonlu durum yapısı yolun Türkçenin biçim dizimince izinli olup olmadığını sınar. Sık fakat izinsiz yol kabul edilmez; seyrek fakat izinli yol yalnız olasılığı düşük diye silinmez.
- Transformer bağlam ayrıştırıcısıdır. İzinli adayları Türkçeye özgü sonlu durumlu çözümleyici üretir; Transformer bu adayları cümle bağlamına göre sıralar. Model tokenları biçim birimi sayılmaz ve açıklamalar kök, ek, bağlayıcı ünlü ile zamir n’si sınırlarına geri hizalanmadan dil bilgisel kanıt olarak sunulmaz.
- Eşyazımlılık adayı gerçek eşyazımlılıkla özdeş değildir. Aynı yüzeyin birden çok lemaya bağlanması gerçek anlam veya görev ayrılığından, özel addan, yanlış lematizasyondan ya da OCR gürültüsünden doğabilir. Canlı kart çoklu lema yüzeyini, geçiş ağırlığını, baskın aday payını ve bağlam çözümleme adayını ayrı paydalarla verir. Bağlam belirsizliği adayı oranı, cümle işlendiği hâlde çözülemeyen token oranı diye sunulmaz.
- Korumalı küme belirli bir hesap için yazım, karakter, sıklık, lema bağı, bağlam ve kaynak kalite kapılarından geçen yüksek kesinlikli alt kümedir. Korumalı sözü kaydın kesin doğru olduğunu veya dışarıdaki kayıtların silindiğini göstermez. Her ölçüm kendi paydasını, eşiklerini ve dışlama nedenini ayrıca bildirir.
- Ek Üretkenlik İndeksi (EÜİ) yalnız ek sıklığı değildir. Gerçekleşmiş farklı lema çeşitliliği, sözlük doğrulamalı tek geçişli türlerin klasik V₁/N oranı, birden çok resource groupnda tanıklanan düşük sıklıklı türetim cephesi, kaynaklar arası yayılım, normalleştirilmiş kaynak entropisi ve kalite kapısından sağ çıkma oranı birlikte hesaplanır. Genişleme sinyalinde klasik olası üretkenlik yüzde 70, türetim cephesi yüzde 30 ağırlık taşır; bu sinyal ile diğer dört ana bileşenin eşit ağırlıklı geometrik ortalaması sıfır ile yüz arasına taşınır. Değer evrensel bir sabit değil, aynı anda ve aynı kapılarla ölçülen ek aileleri arasındaki göreli indekstir.
- EÜİ canlıdır. Kart görünür olduğunda güncel derlemdeki lema ve kaynak tabloları sorgu önbelleği kapalı olarak yeniden okunur; derlem sıra numarası ve hesaplama zamanı sonuçla birlikte gösterilir. Sözlük dışındaki tek geçişli biçimler, olumlu öğrenme etiketi bulunmayan adaylar, etkin ret kayıtları, geçersiz tabanlar, büyük harf ağırlığı yüksek kayıtlar ve Türkçe ses yapısıyla uyuşmayan çözümlemeler skoru büyütmez. Çekimle karışma riski yüksek -cı/-ci and -lı/-li ailelerinin sözlük dışı adayları türetim cephesine alınmaz.
- Türetim cephesi sözlük kararı değildir. Gerçek bir GTS tabanına bağlanan, düşük sıklıklı ve en az iki resource groupnda görülen sözlük dışı biçimleri bağlam incelemesine taşır. Sözlükbirimleşme, anlam kararlılığı, doğru lema, yazım ve cümle bağlamı insan denetimiyle doğrulanmadan bu adaylar yeni madde olarak yayımlanmaz.
- Canlı Türkçe Ek Atlası, etiketli biçimbilim kaynağında sınırı tanıklanmış bütün ek yüzeylerini aynı yüzey-lema çiftindeki çözümleme paylarını koruyarak güncel derlem geçişleriyle birleştirir. Yapım, çatı ve çekim aileleri; ses uyumuna bağlı yüzey değişkeleri, farklı lema ve biçim sayısı, milyon sözcük başına yaygınlık, söz türü, yapı, özellik dizisi, zincirdeki konum ve komşu eklerle birlikte ölçülür. Etiketli kaynağın dışında kalan bir harf sonu yalnız sık olduğu için ek sayılmaz.
- Ek sınırı kaynak denetimlidir. Süer Eker’in Çağdaş Türk Dili kitabındaki biçim bilgisi bölümü, mevcut GTS ek adlarını değiştiren veya bütün sonları otomatik söken bir liste olarak kullanılmaz. Kaynaktaki karşıt örnekler sınır doğrulamasına çevrilir: gör-ü-ş- işteş çatısında dar ünlü ayrı, gör-üş eylem adında -üş tek parçadır; art arda çatı ekleri ayrı ayrı korunur.
- Yardımcı sesler kör kuralla ayrılmaz. Yardımcı y yalnız doğrulanmış ek ailesinde ekten ayrılır; -yor içindeki asli y korunur. Üçüncü kişi iyeliği veya zamir tabanından sonra durum ekinden önceki n, zamir n’si olarak ayrılır; üçüncü tekil iyelikteki -sı/-si biçiminin s sesi korunur. Açık özellik veya bağlam bulunmayan çok işlevli yüzey çözümleme kuyruğunda bırakılır.
- Morfotaktik model, kök veya gövdeden sonra gelen ekleri sıralı bir ağ olarak ele alır. Ek aileleri düğüm, tanıklanmış ardışıklıklar yönlü ve ağırlıklı bağdır; P(ej|ei) koşullu geçiş olasılığı yalnız yerel komşuluğu, tam zincir kaydı ise uzun yolu korur. Zincir uzunluğu, sınıf yolu, geçiş entropisi ve yapım → çatı → çekim çekirdek sıra uyumu birlikte raporlanır. Sıklık dil bilgisel doğruluğun tek başına kanıtı değildir.
- Bağlama duyarlı ve çözümlenmemiş ekler gizlenmez. Aynı yüzeyin araç durumu ile isimden fiil türetimi gibi birden çok işleve açıldığı kayıtlar kesin bağlam bulunmadan tek sınıfa zorlanmaz. Ek sınırı tanıklı olduğu hâlde işlev ailesi güvenle belirlenemeyen kayıtlar çözümleme kuyruğunda oranıyla gösterilir. Özellik dizisinin açıkça doğruladığı -makta/-mekte, ek eylem + geniş zaman -dır/-dir ve kip-kişi bileşimleri ayrı ailelere alınır; aynı yüzey çatı bağlamındaysa bu kararla ezilmez.
- İyelik sınırı çözümlemede korunur. Birinci ve ikinci kişi iyeliğinde ünsüzle biten tabandan sonra beliren dar bağlayıcı ünlü kişi ekine katılmaz; ev + -i- + -miz and kitaplar + -ı- + -mız biçiminde ayrı gösterilir. Ünlüyle biten masa + -mız yapısında ayrıca bağlayıcı ünlü yoktur.
- Zamir n’si üçüncü kişi iyeliği ile ardından gelen durum eki arasında ayrı katmandır; kapı + -sı + -n- + -dan çözümlemesinde ne iyelik ekine ne ayrılma ekine katılır. Tarihsel kökenine ilişkin farklı görüşler bulunduğu için eş zamanlı sınır ile köken iddiası birbirine karıştırılmaz.
- Bilgi paketleme bölümü, bir cümledeki gözlenebilir yapısal yükün çekimli yüklem adayına ulaşıncaya kadar ve yüklemden sonra nasıl dağıldığını inceler. Tamlama ilişkilerini düşündüren ilgi-iyelik işaretleri, fiilimsi katmanları, durum işaretli katılanlar, yüklemin göreli konumu ve kaynaklar arası yayılım aynı canlı pencerede ölçülür; konu, odak ve vurgu gibi söylem değerleri ise yalnız biçimsel ipucundan kesinleştirilmez.
- Bilgi Paketleme İndeksi (BPİ), ön yük yoğunluğu, yüklem ufku, iç katmanlaşma, konumsal esneklik ve kaynak dayanıklılığının açıklanabilir ağırlıklı birleşimidir. Kaynak grupları dengeli tavanla örneklenir; 5–40 yüzey birimli ve tam olarak bir çekimli yüklem adayı bulunan cümleler ana hesaba girer. Sonuç anlam miktarı, anlama güçlüğü, dil bilgisel doğruluk, yazı kalitesi veya estetik değer puanı değildir.
- BPİ canlı ve önbelleksizdir. Bölüm açıldığında en yeni kaynak dengeli cümle penceresi yeniden okunur; toplam derlem sayıları, uygun cümle oranı, morfoloji kapsamı, altı yüklem bölgesi, yapı sınıfları, uzunluk ve kaynak profilleri aynı zaman damgasıyla yayımlanır. Derlemin bütününde insan denetimli bağımlılık ağacı bulunmadığı için tamlama ve öge sonuçları kesin bağımlılık etiketi değil, açık morfolojik işaretlere dayalı çözümleme adaylarıdır.
- Tanıklar eksiksizlik, ilgililik, kaynak kimliği ve nötrlük bakımından insan denetiminden geçer. İdeolojik, cinsiyetçi, dinî tartışmalı, şiddet içeren veya kişileri hedef alan cümleler kullanılmaz.
- Grafikler, özgün şemalar, terim açıklamaları ve doğal akademik dille yazılmış soru-yanıtlar ana anlatımı destekler.
- Toplu içgörü raporu ayrı ölçümleri güncel bir sentezde birleştirir. Canlı derlem toplamları her açılışta yenilenir; ağır uyum, çeşitlilik, nöbetleşme ve entropi ölçümleri sürüm anlık görüntüsüne bağlı kalır. Rapor zaman zaman aynı kalıcı adreste yeniden hesaplanır ve değişiklik bölüm sürüm kaydına yazılır.
- Her bölüm tek kalıcı adreste sürümlenir; ayrı sürüm sayfaları açılmaz. Güncelleme zamanı Türkiye saatiyle, değişiklik özeti bölüm içinde gösterilir.
İlk bölüm: Biçim bilgisi; kök ve gövde, çekim ve türetim, çokluk ve durum, iyelik, ses uyumu, gövde sonu nöbetleşmeleri, eylem katmanları ve sözlükbirimleşme başlıklarını canlı derlem ölçümleriyle ele alır. Kuram ve hesaplama kartları her yöntemin Türkçe veri temsilini, formülünü, gerçek tanık bağlantısını, güncel içgörüsünü, doğrulama yolunu ve yorum sınırını birlikte gösterir. Toplu içgörü raporu bu ayrı bulguları sürümlü bir genel sonuçta birleştirir.
▸Entry History and Academic Citation /madde/<madde>
Each dictionary entry has a persistent content revision identifier, verifiable first-record and last-editorial-revision dates, definition changes, and academic citation text tied to the current content.
- The revision identifier combines the entry ID with a SHA-256 digest of the current headword, definitions, and labels; it changes whenever the content changes.
- The event timeline is projected from the GTS editorial log; user names, file paths, processing batches, and other internal details are not made public.
- When the original publication date of a legacy entry cannot be verified, no artificial date is generated; the digital record date and documented revisions are shown separately.
- APA, MLA, Chicago, and ISO 690 formats include the current revision identifier, entry URL, and access date.
▸Questions and Answers /sorular-ve-yanitlar
GTS maddelerini ve T-BDLD'nin canlı ölçümlerini soru biçiminde açıklayan kamuya açık bilgi alanıdır. Headwords and Meanings / Definitions ayrı süzülebilir; Newly Added to GTS sözcükler de yayımlanmış tanımlanabilirlik kanıtlarıyla kapsama alınır.
- Madde, anlam, tanım, gönderme, köken ve (PHP 3, PHP 4) bilgileri GTS kimlikleri üzerinden doğrulanır.
- Derlem sıklığı, temiz bağlam, alan ve tür yayılımı, eş kullanım, serbestlik derecesi, değerlik, hibrit sözcük ağı ve yaşam eğrisi gerektiğinde birlikte taranır.
- Soru, kısa yanıt ve ayrıntılı açıklama editoryal kayıttır; mekanik kalıp üretim kullanılmaz.
- Geçiş sayısı ve derlem büyüklüğü sayfa açıldığında canlı ClickHouse verisinden yenilenir; açıklama içindeki kavramsal karar değişken sayaçlara bağlanmaz.
- Bağlamlarda kaynak adı gösterilmez; yalnız kamu kalite süzgecinden geçen, soruyu açıklamaya yarayan sınırlı tanıklar kullanılır.
Comment limit: Derlemde gözlenen yaygınlık, alan dağılımı veya tamlayıcı örüntüsü tek başına değişmez bir dil bilgisi kuralı değildir. Her yanıtta kapsam ile sınırlılık bu nedenle ayrıca belirtilir.
▸Corpus Statistics /derlem-alan-tur-istatistikleri
Text corpus (T-BDLD) is a contextual field and text type analysis of Turkish contextual Lemmatization Corpus (Science Art · Technology · Religion · Politics · Economic · Roman poetry · et Vd.): Page 7 is composed of tabs:
- Brief & Guide: Methods and summary numbers. The metric names in the top statistical cards are marked, values are red and aligned.
- Corpus Statistics: Weights, hapax and distinguishing points for each field/structure.
- Word Search: In what fields/types does a word go through, what trust level.
- Area/Structure: Word overlap between fields/types.
- Additional FindingsHapax, no-species word, long tail.
- Findings and Comments: The dictionaryological reading of numbers.
- 📏 Sentence Similarity (geometric / TF-IDF cosine): the exact ratio of the median sentence length, the type-tocene ratio (TTR), the average / median / 90th and 99th percent cosine similarities in row pairs via sample. "Similar / medium similar / less similar" reveals the close neighbouring distribution of corpus re-substantiated/shablonic content.
▸New-Meaning Discoveries /new-meanings
Two types of findings are produced by comparing T-BDLD corpus to GTS:
- New meaning candidates: GTS has a dictionary headword, but corpus is used in a new context (or. dil → programlama dili, şok ▪ Economic tremors).
- New term/dictionary entry candidates: Never on GTS, corpus located word (or. dezenflasyon, algoritma, yapay zekâ).
- Excluded trials: The records produced in the first scan but were not listed because of the fact that it was already on GTS or due to lack of evidence (show for direction transparency).
For each card: the proposed definition draft, existing meanings in GTS, the corpus evidence (evidence, evidence sentence), trigger word, co-existions, sample sentences and location in the field/structure analysis. The definitions are prepared not as a pattern but as meaning explanations in accordance with modern vocabulary criteria; the final decision is the editor's.
▸Corpus-GTS Comparison /gts-varolanlar
corpus passed by and also found in GTS Word usage analysis. Two columns for each word:
- Left (GTS): The meaning of the word in the dictionary, the property labels
- Right (corpus): frequency in the collection, surface shapes, context
Filtering, frequency, or context index, according to the word type, can be arranged.
▸Morphological Family and Derivational Network /morfoloji
The word derivative trees, the production plant productivity and the gravitational distribution. Five tabs.:
- Family Search: When you enter a word, the entire relative word shows the tree structure (▪ eye-eyed glasses ί Glasser). Source filter with (All / Lonely GTS / Lonely New Word / Lonely corpus) a membership can be made.
- Production Additive Production: How many different roots do each production crop connect to? The efficiency of the additions, such as "such as "ticks, cy, --li.".
- Gravity Addendums: The pull-ins By category (Ad State Additives, Iyelik, Person, Plural, Time, Kip, Filily, Other) distribution. Total use for each category is presented with different words, additional formats and sample words. search box And the angels will say: "O my people! Source filter and category filter It can be used together.
- Morphological Neighborhood:A relative (parent, brother, derivative) from a word's name is a segregation step away from its name (parents, brothers, derivatives). The resource filter can select the group. Also, the exponential transcript of the wanted word (which often takes the exponents) is added to the panel below.
- Troubled Species: The candidates for false profiling, including the ghost root or the violation of famous harmony.
Production and so forth. Production supplements produce new words (--ness, --i, -l.); exponentials determine grammatical role (- They, uh, they, they -- they, uh --The system analyzes these two groups separately.
Source filters (Family, Neighbourhood, Gravity Tabs):
• GTS Lemmas with dictionary entry in GTS dictionary
• New word Lemmas bearing the YS candidate mark
• Only corpus Corpus lemmmas (not yet classified) neither
▸T-BDLD Corpus — Contextual Domain and Genre Distributions /derlem-tur-dagilimi
What are every word in the company? contextual field or text type Shows how close you stand. primary field/ mode, Trust level, field/tur entropy and by area/species dominance / PPMI / points The three of them are presented.
21 bağlamsal alan: Bilim, Sanat, Teknoloji, Din, Siyaset, Ekonomi, Hukuk, Eğitim, Sağlık, Spor, Tarih, Coğrafya/Doğa, Felsefe, Dil/Dilbilim, Edebiyat, Toplum/Sosyoloji, İletişim/Medya, Psikoloji, Çevre/İklim, Tarım/Gıda ve Mimarlık/Kent. 3 metin türü (Roman, Öykü, Şiir) ayrı bir eksende hesaplanır.
Seven tabs:
- Brief & Guide: General statistics (massembly sentence, TTR, average sentence length, length percentages), special name filter report and method description. The top summary cards show dark names, numerical values red and aligned.
- Corpus Statistics: Bağlamsal alanların ve metin türlerinin ayrı gruplarda kart görünümü. Karta tıklanınca ilgili sınıfın Seed word, Trust level distribution and lists the high 500 words (search + pages).
- Word Search: 6.003.511 Search between the word; field/structure filter, trust level filter and ranking options. When clicked on the word, scores in all fields/ types are shown by bar graphics.
- Area/Structure: Aynı cümlede birden fazla sınıfın kanıtlarının birlikte geçtiği durumların güncel taksonomiye göre üretilen ısı haritası ve sıralı tablosu. Eşik ayarlanabilir.
- Additional Findings: The most frequent dictionary formats, the longest forms, the most powerful connection to the single field/form, the word, the space/species disorganized word (high entropy) and the most frequent area/type combinations.
- Findings and Comments: The short reading of numerical results in terms of linguistic and corpus quality.
- Sentence Similarity: Fast quality indicators on TF-IDF + cosine, full repeat, close repeat and patternic content level.
Important limitation: Area/type labels are not the ontological class of the dictionary, they are in the collection. contextual usage field or text type. Example. "letter" is not an ontologically legal word, even though it is often passed in the context of the law. This context should be considered while the data is being interpreted.
Custom name filter: About 2.7 million occurence (the 15 percent of the word transition) is excluded from this analysis because it is seen as a capital and abbreviated, so that the recipients are calculated by common words without the pollution of private names.
Method update: Tohum biçimler Türkçe karakterleri korunmuş tam lemma ve multiword eşleşmelerle aranır; ClickHouse'taki yönetici lemma düzeltmeleri tarama sırasında uygulanır ve kesin OCR/gürültü kararları dağılım söz varlığından çıkarılır. TF-IDF benzeri ağırlık, PPMI, birlikte geçiş sayısı, baskınlık ve güven kapıları birlikte değerlendirilir. Search resultsndaki derlem yaygınlık kartları da aynı güncel alan/metin türü dağılımı verisini kullanır.
Destek tohumu: “dönem”, “eser” ve “kaynak” gibi çok anlamlı sözcükler tek başına alan etiketi üretmez. Bir cümle ancak bağımsız ve yeterli ağırlığa sahip başka kanıtlarla birlikte sınıflandırılır. Böylece özellikle Tarih, Bilim ve Teknoloji gibi alanların genel sözcükler nedeniyle yapay olarak büyümesi önlenir.
▸T-BDLD Corpus Dictionary (management console ▪ development stage)
corpus (T-BDLD) Free from GTS, the dictionary is a research dictionary based on dictionary principles. The text is removed from the corpus contexts; the gravitational surface forms do not perform dictionarily entry.
Two chapters.:
- Defined dictionary entry: There's enough vocabulary-science, lemmmas defined by the dictionary.
- Candidate Lemmas: The dictionary headword format is kept as a collection of evidence (which is not considered safe for dictionary lemmma or is likely to be attracted/foreign/partitional format).
Page structure:
- General Statistics panel: total unique format, defined/aday ratio, production date and draft production information.
- Dispersion cards: definition trust (high/ middle/low), definition type, word type, domain.
- Filtreler: Search + promise type + field + definition trust + frequency band + sequence.
- Defined dictionary entry list: The detail + when you click on the line editor's note The entrance opens.
- JSON Export: The button in the upper corner is downloaded into a combined JSON file of all definitions + candidates + editor notes.
Information given in defined dictionary enter detail:
- Descriptions: meaning no, definition text, definition format, definition type, definition trust and production information.
- Example sentences: It's neutral dictionary samples from the security filter.
- Frequent: Total occurrence, frequency, relative ratio (in million), frequency band.
- Domain/Genre: The primary field/substantiation (structure/strength/scene) is a level of trust, domain/genre entropy, scoreboard.
- Coordinates: The most frequent neighbors in the left/right context (pill format).
- Editor Note: Free text note for dictionary entry. When saved, the user name and date are automatically added; in the line With Notes He'll show up with his badge.
▸Reports /reports
Recapulates the comparison between corpus (T-BDLD) and GTS in two current reports: New Word Report and New Meaning Exploration Report.
- The word of the new word candidates is a one-word structure, area, frequency, meaning and context distribution.
- The category, level of trust, type of decision of new meaning and new term candidates, and the distribution of evidence occurrence.
- Parts of the information are: it explains how to interpret the proportions in the graphics according to the current data.
- Rule-based insights: when corpus is reanalyzed, candidate numbers, dominant areas and comments are also renewed.
▸GTS Editorial Workflows /yeni-eklenenler
The last recording of the editorship process and new word additions to the GTS dictionary will be on top. The page is not only a processing log, but also a control area that makes the grounds for entering the new word dictionary visible.
- New word & Add YZ Description It appears with a label; technical admission records are not reflected into the user as a separate type of operation.
- The cards are given dictionary headword, word type, subject/field label, definition, domain and time stamp together.
- corpus evidence The card comes from T-BDLD: a total number of occurences, different number of sentences, proportions per million, prevalence bands, row/center and shapes that come forward.
- Corpus evidence is calculated on the card base after the list is loaded so as not to slow down the page; progress bar appears in the calculation.
- The subject labels are derived from dictionary enterry name and definition hints; not the normative field class. Call and address The label must be used only in the form of open call.
- In standard add-on flow, the word type, field tag, corpus evidence, definition source and new word classifier are both confirmed with positive label.
4. Core Concepts
Lemma — The Dictionary Form of a Word
Lemma (In Turkish)lema" (also referred to as "a word that represents different forms of attraction and derivatives of a word) dictionary format, that is, dictionary headwordFor example, book, books, bookshelf, books. all forms such as "kitap"with his lemme; geldim, gelirim, gelecek, gelenler, gelemezdik all forms such as "gelmek"is represented by lemma.
The Difference Between a Lemma and a Surface Form
- surface format (token): The exact phrase in the text (or. From my book).
- Lemma (lema): This form is dictionary headword in the dictionary (or. kitap).
- Root: The smallest meaningful unit (e.g. Glasses ▪ root: göz).
Lemma, unlike the root, protects the derivatives; she only returns the pull-ins; the original unit in dictionary science is Lemma.
Valency — The Complement Structure of a Lexical Unit
Valency describes how many participant or complement positions a lexical unit opens, the grammatical forms through which it selects them, and the usage frames they form together. With verbs, these may include subjects and types of objects; with nouns and adjectives, they may include case-marked, postpositional, or clausal complements.
What Is Analysed?
- Belirtme, yönelme, bulunma, çıkma, vasıta, ilgi ve eşitlik durumlu tamlayıcılar.
- için, gibi, göre, rağmen, hakkında postpositional complements.
- -mak/-mek, -ma/-me, diye and ki clausal complements.
- The external context of single words, phrases, compounds, and compound verbs treated as complete units.
How Is It Interpreted?
- Core tendency is a government pattern observed frequently and consistently across different documents.
- Strong tendency indicates clear complement selection that does not reach the core threshold.
- Optional and rare mark realisations with lower distribution.
- Circumstantial use marks an element that may express time, place, or instrument but cannot be shown to be governed by the lexical unit.
Live Computation Algorithm
- The surface and lemma sequences of the queried single- or multiword unit are aligned in ClickHouse.
- A multiword unit is treated as one target span; its internal words do not count as complement evidence.
- Overt case suffixes, postpositions, and clausal forms near the target are extracted; sentence boundaries and intervening predicates prevent false attachment.
- Each observation is weighted by distance, direction, document and source diversity, realisation diversity, and the Wilson lower confidence bound.
- Valency frames are built from complements occurring together; Shannon entropy measures frame diversity, while the stability index reflects pattern regularity.
- Evidence volume and source distribution determine the confidence score; core and strongly governed slots determine the estimated valency range.
Because Turkish frequently omits an overt subject, the base valency of a verb includes a nominative or implicit subject. However, core and strong labels are lexical government tendencies observed in the live corpus, not absolute grammatical requirements. When evidence is insufficient, the system does not invent a number and reports no clear government. When new sentences are written to the corpus, the live ClickHouse signature changes and the profile is refreshed without waiting for a manual analysis.
Lemmatization — Converting a Word to Its Dictionary Form
Lemmatization (lemmatization) Downgrade to dictionary format (lemma) It's the process. "I took my books and read them" lematized sentence "to buy and read books" It becomes form. This process means the cast of the pull attachments, the dissolution of irregular forms, and the combination of the word fundamentally.
Lemmatization vs. Stemming
Lematizasyon, dilbilimsel bilgi By using the correct dictionary format. Hullation is a rule-based cut, and often does not produce a real word ("Our books" ▪ hullation: "kitapla"; lematizasyon: "kitap"Lemmatization is both more precise and slower.
The Importance and Challenges of Lemmatization in Turkish
Turkish, an agglutinative (suffixing) languageA word root can take on tens, sometimes hundreds of different forms with the additions of gravity and derivatives added over and over. Lemmatization is relatively simple in less-gravitated languages such as English (cats → cat, running → run); but Turkish, lematization hard and critical It has certain properties:
Sources of Difficulty in Turkish
- Formation stack: Single word can get 10+ additional (As if we can't make it European.).
- Famous harmony and famous/unknown changes: Additional formats vary in sound (-lar/-ler, -dan/-den, -da/-de/-ta/-te).
- Unknown softness: book ▪ book, tree ▪ tree.
- The famous fall: burun → burnu, son ▪ son.
- Foreign word: hukuk → hukuku (It will be said): "Nay! Answer ▪ (softly soft) is irregular.
- United verbs and co-act structures: Help, decide, get lost The relationship between the pieces is subjective.
- Undiagnosed formats: gül The word can be both a name (flower) and a verb (smile); only context determines that.
Why Is It Critical?
- Basic unit in compilational studies: The frequency count, collocation and meaning analysis must be done at the lima level.
- In the dictionary preparation: dictionary enterry heads are held in lemma form.
- Machine translation and NLP: Right lemma means right response production.
- Meaning analysis: Even models such as BERT produce more consistent results with lematization.
- In search engines: "gelmek"calling geldim, geliyor, gelecek He must bring the results.
The Role of T-BDLD
Used in this system corpus (T-BDLD), a continuous growing data set where the correct lemme is solved according to the context of each word. gül (laughs) kaz It means that twin shapes such as animals are separated from the context correctly. This solves one of the biggest lematization challenges for Turkish.
What Would Happen Without Lemmatization?
A computational study without lematization causes serious problems in a language that needs to be added like Turkish:
The first frequency distribution is broken. books, books, bookcases, books. They are like separate words. In fact, the "book" Lemma, which passes 500 times, appears at 30/40 times in dozens of different forms; the word ranking becomes misleading.
The 2nd dictionary comparison collapses. in GTS gelmek has a substance; however, corpus I'm here, I've come, we've arrived If it is not Lematized, these forms cannot be found at GTS and are accidentally marked as the new word.
Colocation III cannot be detected. "to help" and "I helped, they'll help, they helped" It looks like a multiword expression; statistics such as PMI, NPMI, are wrong.
The 4th meaning analysis is off-limits. Even BERT-based vectorization learns the noise of gravity when it's not lymatized; similar meaningful word vectors go away.
The new word discovery makes no sense. In a text that's not Lemmatized From what we can technologicalize. The form looks like a new word, whereas this, teknolojik It's the attractive form of your adjective.
Call and accessibility fail. User "kitap" It calls, but the system only finds this exact format; "the book, the books, the bookcase" It can't bring back the results.
The 7th Academy studies lose credibility. All fields, such as word presence measurements, frequent studies, lexicon design, produce misleading data.
Lemmatization, compilational and dictionaryological studies in additional languages essential Preconditions. All analysis of this system (new word exploration, new meaning discovery, collocation analysis, coverage rate), is based on T-BDLD's contextual lematization output.
How New Words Are Found and Defined
In GTS, a new word is a single- or multiword lexical-unit candidate that has natural usage evidence in the live corpus but is not represented among current entries, entry-internal compounds, fixed expressions, parenthetical alternatives, or established spelling variants. Here, new means new to GTS as an independent entry; it does not necessarily mean that the form was coined recently.
Method transparency and protection: Candidate selection uses multilayer quality filters. This guide explains their purpose and place in the editorial process, but internal rule lists, exact thresholds, and implementation details that could be misused to generate noisy candidates are not published. Frequency or a productive suffix alone never causes a candidate to be entered automatically.
Candidate discovery workflow
- Live evidence: The accumulated candidate pool and all subsequently added corpus sentences are examined together, allowing forms with previously insufficient evidence to be reconsidered when new attestations arrive.
- Separate paths: Single words and possible phrases, compounds, idiomatic structures, or compound verbs are ranked separately because the same decision logic does not fit both groups.
- Surface-form and lemma alignment: Surface forms are aligned with context-sensitive lemma analyses. The goal is to identify the lexical unit realized in the sentence, not merely to reduce a form to a root.
- Quality screening: Automated quality filters narrow the candidate universe. Screening is not an acceptance decision; it only prioritizes records for editorial examination.
- Dictionary collision check: Current entries, parenthetical forms, compounds, idioms, cross-references, spelling variants, earlier rejections, and withdrawn records are checked together. An existing unit is never counted as a new entry.
- Corpus evidence collection: Sentences, document and source dispersion, first and last observation, surface variants, grammatical behavior, and nearby contexts are collected. Duplicate copies of a text do not count as independent evidence.
- Single-word review: Context determines whether a form is merely inflected or functions as an independent lexical unit with a stable meaning. Productive Turkish suffixation may support the evidence but is never sufficient by itself.
- Multiword review: Association strength, stable ordering, semantic unity, transferability across sentences, and the function of naming a recognizable concept are considered together. Freely generated word sequences are not treated as lexical units.
- Establishment and definability: Dispersion across documents and sources is evaluated separately from the ability to derive a portable meaning core from natural attestations.
- Editorial decision: Strongly evidenced candidates proceed to publication, insufficiently evidenced forms remain under observation, and non-lexical records are rejected. Weak candidates are never added merely to meet a numerical target.
What does the evidence package contain?
- Independent, natural context sentences
- Document, source, domain, genre, and time dispersion
- Lemma, surface-form, and part-of-speech analyses
- Derivational evidence for single words and association evidence for multiword units
- Checks against GTS entries, spelling variants, and previous editorial decisions
- A shared meaning core, superordinate concept, and distinguishing properties
How is the evidence interpreted?
- Frequency supports the existence of usage but does not prove lexical-unit status.
- Source diversity is stronger evidence than repetitions of the same text.
- Contextual diversity supports portability, while a stable shared core supports definability.
- Low frequency does not cause automatic rejection, but it limits claims about establishment.
- A domain label describes the corpus environment in which a use is concentrated, not the essence of the word.
- Evidence cards can be recomputed as the live corpus grows.
Definition drafting and verification
- Sense separation: All attestations are read together, while distinct meanings of the same form are kept apart.
- Grammatical function: Part of speech, usage function, and the boundary of a multiword unit are determined from sentences.
- Meaning core: The recurring meaning, its superordinate concept, and properties that distinguish it from neighboring units are identified.
- Original draft: Definitions are written from corpus evidence rather than copied from another dictionary. They do not repeat the headword, rely on mechanical formulas, or become needlessly short.
- Decontextualization: Details tied to one person, place, event, or sentence are removed so that the definition applies to other natural uses of the same sense.
- Three-pass verification: Evidence extraction, definition construction, and validation against the attestations are conducted separately. Unsupported properties are removed.
- Lexicographic review: Clarity, scope, non-circularity, substitutability, neutrality, spelling, punctuation, and agreement with the headword are checked.
- Originality and provenance: Similarity within the dictionary, independence from evidence sentences, the producing model, editorial package, and Turkey-time processing date are recorded.
- Guarded publication: Existing-entry checks are repeated immediately before publication. Approved records enter the current dictionary version, and their definability and live-corpus evidence become visible on public pages.
What happens after publication?
The definition is versioned as dictionary content, while its addition time, producing model, editorial action, and evidence package remain traceable. The definition is stable dictionary content, whereas frequency, last observation, context dynamics, domain and genre distribution, valency, and graph relations are read from the growing live corpus. Later evidence can lead to revision or withdrawal, and earlier rejection or withdrawal records prevent accidental repetition.
New-Word Geometric Similarity (context-vector cosine similarity)
In the Management Console New Word Check The panel opened when a word is clicked on a tab with the relevant candidate Other candidate used similar word are listed as badges.
Method (Does Not Require BERT or Semantic Vectors)
- The context vector: For each candidate, a word count vector on the left side of the corpus/sa is created (most frequently 200 neighbours).
- L2 normalize: Each vector is reduced to the unit length; the frequency difference does not distort the result.
- Cosine resemblance: Between the target vector and the other candidates' vectors, the cosine is calculated; the highest K candidate is rotated.
- Filtre: Candidates without at least three different neighbors are skipped (signal unreliable).
Pratikte: brendi clicked Whiskey, raki, vodka, cognac gibi Same synthic positions. Candidates appear at the top. The percentage shown on the badge is the cosine equivalent ratio (80,, in the same context, 30% of partial overlap).
Calculation is fast (~10-50 ms every interrogation); when the first request comes, all candidate lemons are automatically indexed within the lysical in ~1-3 seconds, and the next interrogations are instant.
Contextualization Index (CI) 0 – 100
In the collection of a word context wealth The measurement is a compound point. It's calculated from these three components:
- Neighbourhood variety: Different word number seen next to the word
- Dispersal balance (entropy): How balanced the use distribution of neighbors is
- Use support: Spreading between the first and last line of the word
Result three banda separate:
The high BI word is used in a wide variety of contexts, a mature candidate for dictionaryization.
PMI, NPMI, LogDice, and Sentence Frequency — Association Measures
Two or more words Is it random or patterned They measure it together. It's used in multiword expression detection and new word conjugations.
PMI (Pointwise Mutual Information ▪ Mutual Information)
The probability of two words being seen together is the logarithm of the ratio of the probability of being seen as independent to multiplication.
PMI = log₂( P(X,Y) / (P(X) · P(Y)) )
The high PMI ί patterned multiword expression.
NPMI (Normalized Pointwise Mutual Information ▪ Normalized Point Controversial Information)
The normalized format of PMI to -1 to +1. It makes it easier to compare and interpret different frequency co-existences.
Close to +1 is the perfect connection. 0..ί independence. Average.86.
LogDice (Logaritmic Dice Coefficient)
It's a measure that's resistant to frequency, and rare and often the word compares fairly.
The theoretical upper limit is 14. >10 ▪ very strong connection.
Sentence Frequency
It shows how many different sentences a co-exist from. It helps separate a pattern from a very repeating pattern in the same text piece that lives in different contexts.
High frequency of sentences can be used across corpus, not a single recursion.
How to Read New Word Co-conspirators Score?
The system is in new word co-existences. NPMI (Normalized Pointwise Mutual Information ▪ Normalized Point Reciprocal Information) + LogDice + sentence frequency NPMI uses the combination. It shows how far apart the union is from coincidence, how stable the union can be against frequency, and how often the sentence frequency spreads to how many different sentences. Multi-worded units at GTS are directly strong evidence; If those not at GTS pass through statistical thresholds The most meaningful co-existences It's in the section, otherwise Unsubstantiated partners yet. Watched in the section.
Derivational Family
Same The entire word that originated from the root. It's the morphological structure that shows together. göz → Eyes → Glasses → Glassesman → Glassing.
For each family member, one of three conditions is determined:
- GTS ▪ word is available in GTS dictionary
- YS ▪ New Word Candidate (only corpus, not in the dictionary)
- D (PHP 3, PHP 4)
Suffix Productivity
It's a product of a derivative. How many different roots? It's the measure that can be added. Very productive additions, such as the "t-t" eclair, are derived by thousands of words, while fewer productive attachments such as "buddy" are limited to faces.
A suffix's only corpus Adding to the new roots we see, that crop It's still an active derivative tool. It's one way to measure the virility of language.
Phrase Formation
One word from you, corpus The tendency to create multiword expression with other words. For example "Uncomfortable" word, corpus "indisposed", "inappropriate." If it is systematically passed in multi-word formats, such as these, then this word shows high massing.
How Is It Measured?
A single word's clustering rate is calculated by this formula:
Example: "edici" He's been through it 823 times alone. "disturbing" (45), "kurum edici" (23)... ifthesets are 180, the ratio of the clusters to the sum is 180 180/823 ≈ %22.
Ratio Bands and Colors
Use on the Home Page
The word that piles up on the new word list It's purple and highlighted. It'll show up.
- 🔗N multiword expression number of words
- Equipment rate: %N ▪ Colored badge according to tape
Clicking on the word is below The panel that opens down It appears and is listed as the entire multiword expression chip (sees multiple times and ratio in each chip). Click on Chip, and that multiword expression is written in the search box and applied to the filter. The panel closes when you click the word again.
From the ranking menu "By the rate of countermeasures" The only word that can be replaced by selection is the one that can become the most clustered. This helps quickly find the molding word cores.
Turkish Suffixes: Derivation vs. Inflection
Turkish supplements are divided into two basic groups: Production extensions produces new word; exponentials The word defines the grammatical role of the word. The system classifies and classifys both additional groups with labeled corpus separately.
Derivational Suffixes (Form New Words)
- -ness,ness, -ness -ness (abstract noun): Beauty, kindness, library
- -cı, -ci, -cu, -cü (occupation): Janitor, passenger, bagel maker
- - ♪ Li, li, lu, lu ♪ (possession): smart, sweet, salty
- - You, you - you, you (absence): homeless, thirsty, foolish
- - Brother. - Stone. (partnership): Comrade, colleague, contemporary
- -sal, -sel It's a date: scientific, agricultural, educational
- - Be, be, be (copular verb): To be united, to be beautiful
- -it's, -it is, - it's... (agent): printer, reader, savior
- - ♪,, g, gu, ugh ♪ (action name): Love, instruments, interrogation
- The day, the day, The day (state): "Wounded, tired, absentminded"
Inflectional Suffixes (Grammatical Function)
- Case: enjoining, completing, turning, finding, dating, means, equality (-ce)
- Requisition: m, n, s, - our, - your, - their
- Plural: -lar, -ler
- Zaman: - see, hear, hear (now), -receive, -future
- Kipler: condition, command, demand, necessity.
- Fiilimsi: participles (-an/-acak/-dık), verbal nouns (-mak/-ma/-ış), converbs (-arak/-ıp)
- Olumsuzluk: -me, -ma
- Interrogative particle: -mm, -mi, -mu, -mu
- Copula: -, -.. (report)
- Care --which: evdeki, sabahki
Why Does It Matter?
- New word discovery: The new word is detected with productive attachments such as "er-er, -lash, -blow.".
- Additional productivity analysis: What crop is measured in how many different roots are connected ί active/dead supplementary.
- Lemmatization accuracy: Production is joined by root and becomes lemma; gravity is added.
- Dictionary editing: The word dictionary is the headword produced by the production add-on, not the shapes with the pull-in.
Voice in Turkish Verbs (Causative, Passive, Reflexive, Reciprocal)
A verb roof, a verb Relationship between subject and object It is an additional category. There are four basic roofs in Turkish, and they are all specially added to the verb root. The system auto-classifies the labeled corpus.
Causative
-tır, -tir, -tur, -tür / -dır, -dir, -dur, -dür / -t ▪ Explains that someone is acting against someone else.
Example: yap- → yaptır-mak; oku- → okut-mak.
Passive
- ll, - l, -l, ▪ The action was done by someone else, not by the subject.
Example: yap- → yapıl-mak; I don't... → görül-mek.
Returned (Reflexive)
--in, -un, --un, ▪ The subject of the action is directed at him. (The context is important because it can be the same additions as theEdilgen.)
Example: Wash- → washn-mak; giy- → giyin-mek.
Work (Reciprocal)
- Work, work, uh, uh-uh, uh... ▪ It says that the action was done mutually among many.
Example: I don't... → görüş-mek; Beat- → dövüş-mek.
Not: Some roof attachments require context, for example. -ın hem edilgen (I don't...) Both rotational (giyin-) The system can categorize safely by looking at the position of the crop and its later pull, markings in uncertain cases as "digital or rotational." It doesn't mislabel..
Spelling Variants
There are two types of words:
- Next/ Separate Spelling: pencil ↔ pencil
- Orthographic difference: rasgele ↔ rastgele, correcting spelling ί correcting unsubstantiated (hâlâ ↔ hala)
Morphological Coverage
The ratio of how much of the word in the collection can be identified by the existing morphological analyzer. Low coverage ( < 20%) shows that the system needs an additional layer of strength in the new and derived word capture. Morphological Family The module is part of this layer of empowerment.
Domain/Genre Analysis Metrics (T-BDLD)
T-BDLD Corpus üzerinde her sözcüğün 21 bağlamsal alan ve bunlardan ayrı değerlendirilen 3 metin türüyle olan yakınlığı ölçülür. Kullanılan metrikler:
Dominance
The ratio between 0 and 1. The word appears. all sentences That's the share of the sentences that carry the field/type tag.
dominance, t) = Together,t) / total_clumm_s)
High domination means that the word is almost alone in that area/structure.
PPMI (Positive Pointwise Mutual Information)
How far above the expected sighting of the word and the field/type. Negative values are reduced to zero.
PPMI(s, t) = max(0, log₂[ P(s,t) / (P(s)·P(t)) ])
Even in low-domination, high PPMI may be rare but important to the distinctive word.
Score
Combining occurrence support with PPMI. The primary domain/genre is selected according to this point; the dominant is also used as a security door.
score, t) (PPMIίs, t), · log (united_sumle+1)
Distinguishment and adequate use support are considered together.
Domain/Genre Entropy
Sözcüğün alanlara ya da metin türlerine dağılımının çeşitliliği. 0, tek sınıfta toplanmayı; daha yüksek değerler farklı sınıflara daha dengeli yayılmayı gösterir. Üst sınır, hesaplanan eksendeki güncel sınıf sayısına bağlıdır.
H(s) = −Σ P(t|s) · log₂ P(t|s)
Low entropy = a strong bonding word (lungs, education, prayer); high entropy = general word (one, to be, to say).
Confidence Level
How solid is the primary field/species appointment? The score is considered together by dominant, area/structure difference and the overall word effect:
- high The word is clearly single-space/structure, reliable field/special tag.
- orta It's got a field/type signal, but it's not distinctive; use it with interpretation.
- low The word has multiple spaces/similar strength; the label is only about.
Genre Status
- single species: Word gives a strong signal of almost single species (or. lung ▪ health).
- raid: A kind of obvious step forward, but other species have weak marks.
- mixed: More than one type of strong signal exists (or. education both education and politics).
- genel: No species is connected in a distinctive way (or. bir, olmak, demek).
Warning: Type labels in the dictionary The ontological type No, it's in the collection. its contextual intimacy The letter may often be passed in the context of the law, but ontologically it is not the word of law. Metrics measure this context; do not ignore the actual meaning of the dictionary material when they interpret the results.
Lexicographic Definability
Dictionaryary scientific descriptivity, acting from natural language witnesses to a single or multi-word candidate whose boundaries have been set, is the degree of being represented by the definition of a dictionary that is not subject to a person, event, or word, supported by evidence, uncycled, in certain and in the same sense.
To understand what's being told in a context contextual resolution, the same form-samulation relationship spreads to different documents and sources SettlementThese are not the same measure, even if they are related to the identifiable, so low frequency, even single testimony, does not automatically nullify a candidate; they only limit the evidence of residence.
Four Independent Decision Dimensions
| Kod | Boyut | The question he answered |
|---|---|---|
| A | Reality | Biçim, çok katmanlı kalite filtrelerinden geçen temiz, doğal ve doğru çözümlenmiş bir Türkçe kullanımı mı? |
| L | Birimlik | Is the correct dictionary headword and unit limit set, structure inflected form or free and random multiword expression? |
| T | Identibility | Can witnesses make a reliable, uncycled, portable description? |
| Y | Settlement | To what extent has the use spread to different documents, resources, species and times? |
Algorithm Execution Order
- Candidate and context collection: The ClickHouse collection includes target sentences, documents and source information, and separate, adjacent, short striped, and attracted changes.
- Reality check: Çok katmanlı kalite filtreleri uygulanır; iç kural listeleri ve kesin eşikler kamuya açılmaz.
- Lema and unit border check: The probability of a single word-inflected form is tested for multiword expression integrity, naming function, syntax boundaries and co-existence.
- Independent rating: Reality (
A), birimlik (L), definitionability (T) and integration (Y) is calculated separately; at an early stage, it is not reduced to a single total score. - Three-pass definition production: The first occurrance only extracts evidence and the core of meaning; the second occrence defines from this structure; the third ochurrence reassembles every argument in the definition to the witnesses.
- Description tests: Cycling, independence from context, substitution, mobility, coverage and separation from familiar meanings are controlled. The properties that do not exist are removed from the definition.
- Decision and publication: Noise and non-unit structures are eliminated; missing context candidates are monitored, except under editor's supervision, carrying strong evidence.
accepted_reviewedThe record of its state is public record.
How Is the Definability Score Calculated?
T The score is made up of seven subdimensions. Unknown dimensions are not zero; they are also zero.
kapsama It shows what measurements can really be evaluated by its value.
| Kod | Subdimension | Denetim |
|---|---|---|
| M | Meaning core | Is there a basic meaning or function? |
| G | Top concept/work | Can you determine general class, process, quality or communication function? |
| D | Differentive property | Is it proven to distinguish between close concepts? |
| C | Non-binding from context | Is the meaning preserved when the person, the location, and the one event specific details are removed? |
| I | Resuscitation and mobility | Is the description reciprocal in authentic use and moving to other possible uses? |
| E | Evidence support | Are the necessary meaning components in the description connected to the witnesses? |
| R | Reproduction | Do at least three independent solutions combine in the same superior concept and meaning core? |
T = 0,20M + 0,15G + 0,15D + 0,15C + 0,10I + 0,15E + 0,10R
Karar bantları ve yorum sınırı
- Gerçeklik ve birimlik kanıtı zayıf adaylar ek incelemeye, izlemeye veya elemeye yönlendirilir.
- For strong descriptivity yeterli kanıt kapsamı, tanım iddialarının tanıklara bağlanması ve Independent analysislerin ortak anlam çekirdeğinde birleşmesi gerekir.
- Yerleşiklik desteği zayıf bir kayıt doğrudan gürültü sayılmaz; anlamı açık olsa bile yeni veri gelene kadar izleme katmanında tutulabilir.
- The percentages on the public card are not the probability of accuracy. Each shows a different decision size and the power of the existing evidence.
Definition Originality Score
The score is an explainable 0–100 measure of how independently a current GTS definition is worded relative to other GTS definitions, its corpus evidence, its headword, recurrent definition formulae, and its recorded production trace. It is evidence of in-system originality, not a plagiarism probability or a legal copyright decision covering every external publication.
O = 0.30U + 0.20B + 0.10K + 0.08M + 0.12A + 0.20P
- U, distance within the dictionary: SimHash bands select nearby candidates; the actual character- and word-shingle similarity is then recomputed.
- B, independence from context: the definition is compared with up to eight evidence sentences so that copying a sentence does not appear as original definition writing.
- K, distance from formulae: recurrent n-grams and mechanical definition patterns reduce this component.
- M, independence from the headword: circular wording and direct repetition of the headword are penalized.
- A, structural individuality: explanatory length, content vocabulary, lexical diversity, and sentence structure are evaluated together.
- P, production-trace integrity: the source type, creator, model, Türkiye-time timestamp, raw record, and editorial review are checked.
The displayed formula applies when both corpus context and production trace are available. A missing dimension is not treated as zero. The remaining weights are renormalized to sum to one, and evidence coverage is reported separately. A definition-text hash must match the current active definition; after an edit, the stale score is withheld until recalculation. Redirect-only records are not independent definitions and are therefore marked not applicable.
How the average is calculated
The average shown on the Newly Added page is the unweighted arithmetic mean of every active,
current, applicable definition in the full public-new-entry set:
average = sum of current scores / number of applicable definitions.
Each definition has one equal vote. Pagination, search, and topic filters do not change the
value. Redirects and definitions whose score is pending after a text change are excluded.
Home-Page Tabs and Tools — What Do They Do?
At the bottom of the main page New Entries in GTS, Proverbs, Idioms and Compound Verbs with list tabs 17 solvent tools list tabs show GTS records and T-BDLD prevalence together, while the analyzers use BERT-based word vectors, cosine similarities, clustering and corpus statistics.
1. New Entries in GTS
It is a public list of new dictionary enterry and meaning records added to GTS through editor control.
- How does it work: New dictionary entry, new meaning and reference records are listed by the time they are added; when clicked on the word, the description, badges, and corpus evidence details are opened.
- Question answered: "What new dictionary enterry and meanings have been added to the GTS recently?"
- Output: For each record, the dictionary enterry connection, field/strength badges, additional information and corpus evidence if necessary.
- Proof of identifiability: In evidenced publications, reality, unitity, identifiability, and integration are separated. The card sums up the number of natural witnesses, the agreement on independent analysis, the common meaning core and the definition tests.
How Should Definability Evidence Be Interpreted?
Reality measures the clean and natural use of the form; the dictionary integrity of the unit, one or multiple word structure; the definitionability of natural witnesses can be derived from non-cycled and other uses; and integration measures the spread of use into documents and resources.
The public card is seen in the records that pass the tests of independence, substitution and mobileity through at least three independent resolutions. Percentages are not necessarily a probability of accuracy, but a sign of different decision sizes.
Go to a detailed explanation for the definition of concept, points and decision-making thresholds.
2. Proverbs, Idioms, and Compound Verbs
Shows the stereotyped statements at GTS along with the prevalence of corpus (T-BDLD).
- How does it work: The proverb, phrase and joint verb records are listed alphabetically; the original is listed.
...No nail, no punctuation, no parentheses, doesn't affect the ranking. - corpus connection: Each record shows the number of T-BDLD orbits, proportions per million words and prevalence bands.
- Domain/Genre badges: All domain badges are displayed in descending order by percentage and record count; selecting a badge filters the list to that domain.
- I'm in the proverb/space: Metaphorical patterns don't look at individual words in a single pattern; GTS definitions are also used.
- Cumulative graph: The chart on the list collects the records of the active page or search; the number of captured registrations, total occurence, /M ratio, band distribution, and several of the most frequent patterns appear.
- Comment limit: The list level uses fast-limma/face n-gram matches. Extra scans on the detailed dictionary enter page in very long or flexible patterns can also provide evidence.
Inflected-Form Timestamps
This page opens on the main page with a button shows the first and current frequency of the gravitational surface formats captured by the system in corpus (T-BDLD).
- First capture: If a pair of lemma+inflected forms are first seen, the time stamp will be written.
- Frequent update: If the same format is seen again in later corpus updates, the first time stamp will not change; only the current frequency and the final analysis time will be renewed.
- Start point: 05.04.2026 18:37:22, and prior records do not appear to be one by one ancient history; All corpus data It's in the starting set.
- Old recording start: If there is, the records in the old new word/inflected form history are used. The formats from the starting point show up in the same set of beginnings that are old or incomplete.
- Sorting: The list is reordered according to the first time of capture; the search works both in and out of material.
- Sorting criteria: The most recent captured, oldest captured, frequency high, last analysis time, dictionary headword A-Z and inflected form A-z option can be used.
- Denetim: The forms marked with ocri error, spelling error, meaningless surface, false lema, punctuation residue and similar labels in the management area are written on permanent control register. The rejected formats fall from the user page.
- Technical record: Corpus-scale data is written on permanent paintings at ClickHouse; as analysis is completed, ClickHoose search and analysis layers are updated.
Word Life Curve
This page opens with a button on the main page shows the periodical appearance and resource distribution of a word or part of the T-BDLD.
- Time axis: If the source has year information, the source year is used; if there is no year, the sentence's compilation entry month is used.
- Comment limit: The collection entry month is not the first date of the word's use in Turkish; it only shows the time of entry into the T-BDLD system.
- Single word and multiword expression: With a single word Lemma match, multiword expression is called in a series of consecutive lemma.
- Source distribution: Book/PDF/EPUB, Magazine Park, news, column and web sources are grouped separately.
- Technical record: The timeline and resource distribution are calculated live from the T-BDLD sentences on ClickHouse.
4. Similar Meanings
It finds the closest meaning in the term "BERT" vector space.
- How does it work: The BERT vector for the word "interrogation" returns the first N with the highest score of cosine similarities to all GTS meanings.
- Question answered: "What are the meaningful relatives of this word?"
- Example: joy (with a similar score)
4. Analogy
An analogy completion with vector arithmetic: A → B ise, C → ?
- How does it work: Klasik
B − A + CThe vector process is the closest GTS to the result (with word2vec style but with BERT vectors). - Question answered: "Ankara ί Turkey, Paris ί?" or "king ί queen, male ??"
- Academic value: Proof of how well the model learned semantics.
5. Semantic Field
It removes members of a theme or conceptual field (semantic field).
- How does it work: A few "tohum" words were given (e.g. cat, dog, lionThe average proximity to these seeds in vector space is the other members of the same field (tilki, kurt, leopar…).
- Question answered: "What is the word in Turkish for animal / color / emotion?"
- Use: Theme dictionary preparation, conceptual mapping.
6. 3D Map
It takes 768-dimensional BERT vectors of GTS meanings down to 3 dimensions and visualizes them.
- How does it work: The vectors are downloaded to 3B with PCA or T-SNE ▪ Interactive map that can be rotated, zoomed with Three.js.
- Question answered: "What does the semantic distribution of the dictionary look like in general? What clusters are there?"
- Groups: Each color represents a different set of k-means; the mouse click shows which word/sense that point is.
7. Word Network
Meaningual relationships between the word graf it shows.
- How does it work: K-en close neighbour (kNN) graph ί the most similar k-next of every word connects to it. Network X + D3.js are played with force-directed slideout.
- Question answered: "How is the semantic neighborhood around this word knitted?"
- Yorum: Intense knots are hub words (very connected); isolated knots are semantic lonely words.
8. Semantic Evolution
It shows how a word changes the context of use over time.
- How does it work: The term is divided into corpus occurence periods ί for each term the co-existent vector ί is the "evolution curve" with inter-term cosine differences.
- Question answered: "Virus Is the word used in the 2000s and the 2020s the same way?"
- Limitation: Corpus gives an estimated result when there is no document date information.
9. Word Cloud
The word frequency visualization in the collection is a capitalization.
- How does it work: Frequent meters match the size of the font, ▪ the word creates a cloud with random settlement.
- Question answered: "What word goes most in Turkish?" (visual summary)
- Filtreler: Stopword list + minimum frequency threshold.
10. Semantic Path
It finds the shortest semantic route between two words.
- How does it work: On the word network graph, the Dijkstra algorithm should go from A to B through which nodes should be passed.
- Question answered: "kitap and teknoloji Which word is the semantic bridge between?"
- Example route: book ▪ summer ▪ information ▪ digital ▪ technology
11. Topic Analysis
corpus separates into thematic clusters (Topic Modeling).
- How does it work: LDA or BERT-based clustering ▪ most defining word + example sentences for each set.
- Question answered: "What issues are these corpus talking about?" (automatic theme summary)
- Output: The title for each subject (auto-elected representative word) + key terms + representation sentences.
12. Rare Words
Corpus is a rare (five times) word list.
- How does it work: It filters the lemmmas under the threshold of the frequency counters.
- Question answered: "Which one of those words is the only one-two-time, non-hapax?"
- Use: Spell error detection, expert terming, for the discovery of archaic use.
13. Polysemy
He does a dictionary enterry analysis on GTS, which means more than one thing.
- How does it work: The distance between the meanings of each matter is calculated by BERT vector ▪ the meaning of how different it is seen.
- Question answered: "What word meanings are so far apart?"yüz ▪ organ, number, like the beach)
- Metrik: The average pair of meanings is the distance cosine of the cosine, which is very high = very meaningful, very close meaning.
14. Semantic Diversity
It measures the homogeneousness of the distribution of meaning throughout the dictionary.
- How does it work: The word semantic clustering degree is ί the integer analogy and so on, between clusters.
- Question answered: "How diverse is the word presence in Turkish? What areas are dense, what areas are sparse?"
- Output: The set ratios with the cake graph plus the diversity index (Shannon entropy).
15. Monosemy/Polysemy Comparison
It compares the only meaningful (monosemic) and very meaningful (polisemical) word features.
- How does it work: The only meaningful group, etc., is compared to the frequency, the word type, the field label, the conjugal diversity metrics for the very meaningful group.
- Question answered: "Is it more common for a very meaningful word? What kind of word are they gathering in?"
- Bulgu: The police word usually passes more often (zipfian distribution).
16. Figurative-Language Analysis
It detects possible metaphoric uses.
- How does it work: Large deviations between a word's basic meaning vector at GTS and the utilization vector in the compilation are possible metaphor signals.
- Question answered: "Aslan Is the word "real animal," or is it used in the phrase 'tough person'?"
- Metodoloji: The match (Lakoff-style conceptual metaphor) is the source of the target.
17. Advanced Search
Güncel GTS maddelerini ClickHouse üzerinde maddebaşı, tanım ve editoryal niteliklerle tarayan araştırma aramasıdır.
- Metin eşleştirme: Tam, içerir, başlar, biter, bütün sözcükler, sözcüklerden biri veya RE2 düzenli ifade.
- Editoryal süzgeçler: Köken dili, alan, söz türü, kullanım etiketi, özel ad durumu, tekli/öbek yapı, anlam sayısı ve tanım uzunluğu.
- Example: Maddebaşında
^bil.*lik$düzenli ifadesi ve tek sözcük süzgeci It can be used together. - Output: Paylaşılabilir sorgu bağlantısı, sayfalı sonuç ve CSV/JSON dışa aktarma.
18. Entropy Analysis
The word measures its semantic predictability (Shanon entropy).
- How does it work: The probability distribution of each word in which contexts is the ▪ entropy account.
- Question answered: "Which word is found in many contexts, which are narrowed down?"
- Yorum: High entropy = general word (one, to be, to say); low entropy = special term (lung, Supreme Court).
19. Context Dynamics
It examines the stability and variable of the context of a word.
- How does it work: The vectors of the sentences spoken by the word are ί variants and cluster density ί detection of the "bindset" set.
- Question answered: "Is this word always in the same context or in different areas?"
- Output: Computation diversity score + dominant context clusters.
General Tip
There is an interrogation box at the top of each tab; most tabs trigger a BERT-based account with a "search" or "solut" button. The result graphs are produced by Chart.js, Three.j's or D3.Js. Data source: BERT sentenence-transformer (BERT).emrecan/bert-base-turkish-cased-mean-nli-stsb-tr), GTS dictionary database (99.238 dictionary entry / 133.041 meaning) and T-BDLD corpus (130,29 The processing time depends on tab: simple searches are instant, 3D Map / Word Network / Subject Solution / Few seconds.
5. Glossary of Terms
This is the part of the system. Artificial intelligence, natural language processing, machine learning, corpus linguistics and dictionary science It contains short, descriptive definitions of terms. The term titles correspond to English in Turkish name + parentheses. Under each title, there is a brief example of what concept is, how it is, and how it's used and needed.
A. Artificial Intelligence (AI) Terms
Artificial Intelligence (Artificial Intelligence, AI)
The capacity of computer systems to mimic human intelligence, such as learning, reasoning, understanding of language, and decision-making. A broad umbrella term: learning machines, deep learning, natural language processing and computer vision subspaces.
Narrow Artificial Intelligence (Narrow AI / ANI)
The system that specializes in one task (playing chess, learning sound, translating etc.). All practical applications of today are in this category; insan benzeri genel zekâ yoktur.
Artificial General Intelligence (Artificial General Intelligence, AGI)
The hypothetical system that can perform any cognitive task that one person can do, transfer learning from one field to another, is not yet available; it is the subject of academic goal and debate.
Large Language Model (Large Language Model, LLM)
A neural network model with billions of parameters, trained in broad text compilations. Examples of GPT, Claude, Gemini, LLAMA. Text production, translation, summation, QA do tasks with a single model.
Transformer
It was introduced in 2017, dikkat (attention) It's a neural network architecture based on its mechanism. It works in parallel with the sequenced data (text, sound); it is the basis of all modern large language models, such as BERT, GPT, T5.
Attention Mechanism (Attention)
In a sentence of the model How much you're going to weigh in on what words It's the structure that makes him dynamically learn. "I read the book." okudum kelimesini yorumlarken book' gives high attention to you.
Token
The language model uses text minimum unit. It's not always "word"; sometimes it's subword (subword) or character. computer + education It's divisible by two tokens.
Prompt
The text given to the language model entered ί question, instruction or context. The output of the model depends directly on the content and shape of the prompt.
Prompt Engineering (Prompt Engineering)
The discipline of prompt carefully designed to get the output requested from the language model."You're a dictionary editor..."), it contains techniques such as example, step-by-step thinking.
Fine-Tuning (Fine-tuning)
Re-educating a model trained in a large collection (or. medical texts, legal contracts) with a smaller and more private data set (or medical texts), is much cheaper than training from scratch.
Retrieval-Augmented Generation (Retrieval-Augmented Generation, RAG)
When producing answers to the language model It's from an external database. The RAG on GTS means "responsive and model response.".
Zero-Shot Learning (Zero-shot Learning)
The model is capable of doing business only by instruction, without seeing any special examples for the mission.Translate this sentence into Turkish" When you say, "you do not need to show the illustration of translation.
Few-Shot Learning (Few-shot Learning)
The Prompt's quick capture of the task pattern of two-five examples of offerings is more decisive than that of Zero-shot.
Chain of Thought (Chain-of-Thought, CoT)
The method that allows the model to think about complex questions step-by-step, adding "Take a step" to Prompt, especially in mathematics and logic questions.
Hallucination (Hallucination)
The language model is to produce outputs that do not fit the truth, which are completely fabricated but are presented with confidence. The most important issue of LLMs is the credibility problem; cross-checking is mandatory in dictionary/knowledge studies.
Embedding (Embedding)
The word is the transformation of the sentence or document into a numerical vector (usually in 300/1024 dimensions). It is close to each other in the same meaningful word vector space. All meaningful search, clustering, is the basis of analogy calculations.
Generative Artificial Intelligence (Generative AI)
New text, image, sound or code-making model family. LLMs, video-producing diffusion models (Stable diffusions, DALL·E) and sound models are part of this category.
Multimodal Model
The system can process multiple data types (text + image + sound) in the same model. Models such as GPT-4V, Gemini, can see and produce both images and writing.
B. Natural Language Processing (NLP) Terms
Natural Language Processing (Natural Language Processing, NLP)
The subsidiary field of artificial intelligence, which is interested in the computer understanding, analyzing, and producing human language. Translation, summation, meaning, conversation, word/linguology analysis comes into this area.
Tokenization (Tokenization)
The text can be computerly processed to the smallest units (token) separate. The word-based, sub-word (BPE, WordPiece) or character-based.
Lemmatization (Lemmatization)
Do not download the way a word is drawn into the dictionary format (lemma). From our books → kitap; He came / came / coming → gelmekThe corpus of this site is limazed.
Stemming (Stemming)
The simple form of lematization is that it snaps the attachments roughly. I run. → Run-The dictionary doesn't produce shapes, statistically enough.
Part-of-Speech Tagging (Part-of-Speech Tagging, POS)
Do not name each word, adjective, verb, envelope, pronoun tag. It is one of the first steps in solving the sentence structure.
Named-Entity Recognition (Named Entity Recognition, NER)
Do not identify specific names and categories such as the text, the institution, the location, the date, the amount. Ali went to Ankara yesterday → Ali: Person, Ankara: Location, yesterday: Zaman.
Syntactic Parsing (Parsing / Syntactic Analysis)
There are two approaches, including the subject, the verbum, the complements, the addiction (dependency) and multiword expression-structure (constientuency).
Word Embedding (Word Embedding)
The technique that transforms each word into a constant-dimensional numerical vector, similar meaning to the word takes close vectors, Word2Vec (2013), GloVe, FastText classic examples.
BERT
Google released in 2018 on transformer basis bidirectional The language model, every word in a sentence, Both right and left In context, code. The model family used on this site for GTS meaning vectors is based on BERT.
Word2Vec
In 2013, the word burial technique developed by Google's Tomáš Mikolov and his team was the first successful method to capture word meaning relations with vector arithmetic. "King − male + female ▪ Queen"for example.
Contextual Embedding (Contextual Embedding)
BERT and its successors provide this: "One hundred people" and "He washed his face" in sentences yüz It's represented by different vectors.
Subword Tokenization (Subword Tokenization, BPE / WordPiece)
The word is to shrink the vocabulary by dividing it into common subsections. "of hard work" → Work + blood + lyric + tanIt makes it easier to teach a model of rare word and morphological rich languages.
Language Model (Language Model, LM)
A model that calculates the probability of a word sequence or predicts the next word, according to previous words, from classic n-gram LM to modern transformer LM.
Sentiment Analysis (Sentiment Analysis)
Do not determine whether the text is positive, negative or false. It is common on comment sites, social media analysis.
Machine Translation (Machine Translation, MT)
Do not automatically translate text in one language into another language. Statistical MT (SMT) evolved into (NMT), a national MT ("NMt") (LLM-based MT).
Stop Words (Stop Words)
The word that doesn't make sense, but often passes: And, with, one, thisText analysis is often excluded, but in stylistic studies, it carries information.
Annotation (Annotation)
The process of adding information (such as word, meaning, emotion, private name etc.) to a text by experts. Gold standard This is how education and test data are created.
Inter-Annotator Agreement (Inter-Annotator Agreement, IAA)
The same two/three people who labeled the same data independently have made the same decision. Cohen's Kappa is calculated by measurements such as Fleiss Kappa; >0.80 is considered reliable.
Gold Standard (Gold Standard)
A set of tagged data that has been passed through expert control, considered a reference. The evaluation of models is based on this.
C. Machine Learning Terms
Machine Learning (Machine Learning, ML)
The algorithms should learn from data instead of being clearly programmed. It is divided into three ways as controlled, unsupervised and reinforced learning.
Supervised Learning (Supervised Learning)
Trained with labeled data: each entry has a correct output (ethicet). Classification (sense, species) and regression (numberial estimate) are in this category.
Unsupervised Learning (Unsupervised Learning)
Buildings from non-tag data. Sculpture, dimension reduction, anomaly detection. This is the basic education paradigm of Word2Vec.
Reinforcement Learning (Reinforcement Learning, RL)
The agent learns by interacting with the environment with reward/imposed signals, games, robotics, LLMs. RLHF My name is in this category.
Deep Learning (Deep Learning)
ML subspace based on multilayered artificial neural networks, image recognition, voice processing, revolutionized NLP, LLMs are a product of deep learning.
Overfitting (Overfitting)
The model's ability to memorize the education data and generalize it to new examples, like the student who memorized the book for the exam but couldn't do it when the question changed a little bit.
Underfitting (Underfitting)
The model is so simple that it can't capture patterns in data, it's a low success in education and testing.
Regularization (Regularization)
Techniques (L1, L2, dropout) that add the term punishment to the model to prevent extreme harmony limit model complexity.
Cross-Validation (Cross-Validation)
Don't take the most common 5-fold, 10-fold, by taking each piece into a set of tests in order.
Train/Test/Validation Split (Train/Test/Validation Split)
Do not separate data into three parts: education (the model teacher), verification (the hyperparameter selection), test (the final measurement of success).
Loss Function (Loss Function)
It's the formula that makes the model's predictions digitize how far away from real labels. Education tries to minimize that value.
Gradient Descent (Gradient Descent)
It's a local minimum approach algorithm, which is the engine of all neural network training, taking a step towards the steepest descent of the missing function.
Backpropagation (Backpropagation)
The method that calculates the graph for each weight by repulsing the bug from exit to entry in the neural network, which allows modern depth learning to be possible.
Epoch, Batch, Mini-batch
Epoch: Once the training data is fully circulated. Batch: It's a sample group that takes every step of the way. Mini-batch: Generally, small groups of 32.512 examples.
Learning Rate (Learning Rate)
The coefficient that determines how much weights will change in each step of the Gradyan landing. It explodes very high, very low, and it learns slowly.
Dropout
The technique to randomly disable some of the neurons every time they take a training step. Mussiness katar.
Confusion Matrix (Confusion Matrix)
The table that shows the rating performance: right positive, wrong positive, right negative, wrong negative numbers.
Precision / Recall / F1
Precision (kesinlik): The positive ratio of positive to really positive. Recall (emotion): How much of the real positives did you get? F1: They're both Harmonic Average.
Clustering (Clustering)
K-means (a certain number of clusters) are the main algorithms based on DBSCAN/HDBScan (based on density).
K-means
The most common cluster algorithm. K sets the cluster center, drops each point to the nearest center, updates the centers, repeats them. The number K must be specified in advance.
HDBSCAN
The density-based hierarchical clustering. It determines the number of clusters itself, different labels the noise (isole points). It is preferred in new meaning discovery.
PCA (Principal Component Analysis, Basic Compositive Analysis)
The technique that reduces high-dimensional data to less dimensional linears. It's used for visualization and noise reduction.
t-SNE / UMAP
The high-dimensional data used to visualize 2B/3B non-linear They're dimensional reductions, which visually reveal semantic clusters.
Hyperparameter (Hyperparameter)
The model cannot learn, the value set by hand before training (the speed of learning, the number of layers, the dropout ratio) is a separate specialty.
D. Corpus Linguistics Terms
Corpus (Corpus / Korpus)
A collection of planned and documented text used in linguistic research. It can be verbal or written; it can be explained (annotated) or raw. The data source of this site may be available. T-BDLD derlemidir.
Representativeness (Representativeness)
Corpus is a measure of how much it reflects the alleged language it represents: species, periods, context, resource diversity.
Balance (Balance)
The proportional distribution of different types of text in the collection (the paper, the novel, the academic, the verbal) gives a disorganized corpus bias.
Token / Type
Token: Each word in the text is occurrence (a sentence called "house house" two tokens). Type: Unique word format ("home home" = 1 type).
Type-Token Ratio (Type-Token Ratio, TTR)
The unique word number / total word number. The size of the word variety. It's inversely proportional to the length of text; therefore using normalized variants such as MATTR, STTR.
Hapax Legomena
Derlemde Only once The passing word (Greek "said once") is the natural result of the Zipf law, with the collection of each language ~40-50.
Zipf's Law
"N. most frequent word, 1/N of the most common word." The universal statistical feature of natural languages.bir, olmak), the long tail is numerous but not often.
Collocation (Collocation)
The tendency of two words to go together more than coincidence. Dark tea, strong coffee, making decisions They're co-existent, which is calculated by measurements like PMI, LogDice.
N-gram
N consecutive word sequence. Unigram: One word; bigram: two words ("It's beautiful"); trigram: Three words, the basic unit of co-existence and language model.
Concordance (KWIC) (Concordance, KWIC)
The list of all the contexts in which a word is passed in the middle is the classic tool of dictionary.
PMI (Pointwise Mutual Information, Point Controversial Information)
The logarithmic measure of how much the two events (word A and word B) are seen together compared to coincidence. PMI(A,B) = log₂[ P(A,B) / (P(A)·P(B)) ]Negative values could also come out.
NPMI (Normalized Pointwise Mutual Information)
Normalized Point Reciprocal Information. It's normalized to PMI's [−1, +1] range. +1 = sure don't cross together, 0 = chance, −1 = never being together. It makes the co-existence comparable.
PPMI (Positive PMI)
The measure using only the positive part of PMI. Weak ones are expected to be considered zero; the field/text type distribution is used to show whether the word-alarm/type bond is distinctive.
LogDice
It's called a "strength coed" or a logarithmic variant of the Dice coefficient. Around 14 is considered a "multiple co-existent"; it's more stable in rare bigrams than PMI.
Sentence Frequency
It shows how many different sentences a word or co-existence appears in. In multi-word candidate and co-axis accounts, a single text is used with NPMI and LogDice to separate the pattern of reassembled use.
TF-IDF (Term Frequency × Inverse Document Frequency)
Classical weight that combines how often a word is in a document (TF) and how rare (IDF) is in all documents. Search and document are used in classification.
Cosine Similarity (Cosine Similarity)
It measures how similar the two vectors are in direction. The closer we get to 1, the more context or meaning similarities are used in sentence similarities, word neighborhoods and semantic search modules.
Keyness (Key-ness, Keyword Extraction)
A subcorpus (or. economic news) is a characterizing word extraction based on a reference collection. Log-likehood test or ki-kare.
Diachronic / Synchronic (Diachronic / Synchronic)
Artificial (diachronic): The change of language in time (sympathetic shift, new word). Synchronic (synchronic): The structure during a period (the new meaning of this site is synchronised).
Contextualization Index
The compound indicator (0-100) measures how many contexts a word can be used in this system, and is calculated from the combination of neighboring word diversity, entropy and support (occurrence).
Domain/Genre Entropy
It shows how much a word spreads to contextual areas and types of text. Low entropy is a single area/culture proximity, high entrope is a general word used in many areas/structures.
Dominance (Dominance)
The share of a particular field/species in sentences a word passes. It is one of the main indicators in the field/text distribution that monitors whether the high score is reliable.
Morphological Coverage
How much of the word in the collection can be described by the current morphological analyzer. Low coverage in rich morphology languages such as Turkish requires a layer of strength to capture new word production.
E. Lexicography Terms
Lexicography (Lexicography)
The art and science of making dictionarys. It has a dimension of both practical and theoretical (on word meaning and structure). GTS is a word science product.
Dictionary (Dictionary)
It is an application that lists the words of a language in a specific measure (alphabetic, semantic, subject) and provides definition, origin, sample for each of them.
Headword (Headword / Lemma)
The word format given in the dictionary as the title.gelmek), names are simple (kitap) verilir. GTS'te madde The column holds the dictionary entry head.
Entry (Entry)
A dictionary headword, and all the meanings attached to it, the patterns, the labels, the origins, the proverbs and so forth. A substance on GTS can have multiple meanings.
Meaning (Sense)
It's one of the different concepts that a word carries. yüz ▪ (1) organ, (2) number 100, (3) revolt ("not to face").
Definition (Gloss / Definition)
The short text explaining a meaning. Typically, the sex (kind of...) and distinguishing properties (...olanThe TDK definition tradition is based on the Arististian definition principle.
Polysemy (Polysemy)
The same word related It has different meanings. ayak The organ, the mountain skirt, the deck of cards; they're all conceptually connected.
Homonymy (Homonymy)
The words are the same, but they mean the same. tamamen ilgisiz Word. yüz (organ) yüz (number) is a common-name, historically different roots.
Synonymy (Synonymy)
Two different words have the same or very close meaning. Full co-existence is very rare; it's usually style, frequency, context difference. student / student.
Antonymy (Antonymy)
Converse meaning relationship.hot/ cold), complement (dead/alive) or relational (teacher/teacher) olabilir.
Hyponymy / Hypernymy (Hyponymy / Hypernymy)
Class-low class relationship. Dog The hyponym of "animal" is substantive; hayvan The hyperonomy of "dog." WordNet is built with these relationships.
Meronymy / Holonymy (Meronymy / Holonymy)
Meronim: The piece (mary of the car). Holonim: All of it.
Metaphor (Metaphor)
The meaning of a concept in one field being moved to another. "the footsteps of time", "knowledge stream"It's a key source of new meaning exploration.
Metonymy (Metonymy)
It's the use of something to replace something that's closely related. "Ankara karar verdi" (Ankara = government); "Have a glass" (bardak = inside).
Semantic Field (Semantic Field)
A word cluster that gathers around a particular concept. Food area: Frying, cooking, scalding, cooking... semantic field analysis provides integrity in the dictionary structure.
Usage Label (Usage Label)
Note in which context a meaning goes: In public language, slang, old, formal, metaphoricalIn GTS, the "use" features are in this category.
Domain Label (Domain / Subject Label)
It shows what area of expertise a meaning belongs to: medicine, law, religion, astronomyThe GTS uses 44 canonic fields.
Origin (Etymology)
The historical source of a word and the changes it takes. Telefon ← Yunanca tele (uzak) + phone (ses). GTS'te lisan The column shows the original language.
Loanword (Borrowing / Loanword)
The word taken from another language by a language. In Turkish, Arabic, Persian, French, English quotes are intense. The quote word can eventually adapt sound and meaning.
Neologism (Neologism, New Word)
The new word that just entered the language.bilgisayar), quote (dezenflasyonThe new word module in this system focuses on the search for neology.
Spelling Variant (Orthographic Variant)
Corpus is a standard or preferred form of a format that is seen in GTS. pantalon formats such as these can be kept on the new word list, but in the line, the GTS' response (pantolonThis does not mean that the word is considered a new dictionary entry.
Archaism (Archaism)
The word/form that falls out of language, only remains in ancient texts or official rhetoric. bizatihi, mezkûr, el'anAn archaic dictionary entry detection is an important step for the dictionary update.
Derivation (Derivation)
Adding production to a word and creating a new word. eye ▪ Glasses ί Glassesman ί SpecteringTurkish comes with the richness of derivatives.
Inflection (Inflection)
The additions to the word's linguistic role (position, person, time, number). same It gives different forms of the word. evler, eve, evden → ev.
Compounding (Compounding)
Two or more words combine to form a new meaningful whole. Refrigerator, Monday, honeysuckle. The combined word in Turkish dictionaryology is a controversial area (separation/finished writing).
Proverb (Proverb)
The GTS proverb table contains the saying "16,666.".
Idiom (Idiom)
A phrase that cannot be removed from the meaning of their words, which is used in patterns. Getting mad, putting flour on the rope. is a multiword dictionary entry.
Lemmatization Template
The broad document used in this system defines the rules and exceptions of morphological analysis for each lemma. The router for the production and attractions of Turkish.
F. System and Operations Terms
Local CPU Worker Pool (Local CPU Worker Pool)
Sharing heavy computing jobs among server processors speeds up phases such as corpus reading, candidate collection, multi-word surface formation, context cache preparation and field/text type distribution.
Parallel Processing (Parallel Processing)
Instead of running a job in single file, do not do it in the same time as breaking it into pieces. The large text is used to reduce the waiting period as Corpus grows.
Cache (Cache)
It's often used or calculated that expensive outcomes are kept temporarily or permanently, so that when the same information is asked again, it doesn't recalculate, it opens faster.
Binary Cache (Binary Cache)
The fact that the analysis results are stored in a binary way that can be read quickly by the machine, especially for large-scale tables, surface shapes, and morphological families.
Warm Start (Warm Start)
The approach to fast-tracking the system using pre-made cache and indexes after the service reboots.
Manual Analysis
The purpose of the main analysis process is to be started by the manager after the corpus upload or add-up, so data addition, cleaning and computation times are separated.
Deduplication (Deduplication)
The process of preventing the addition of the same or very similar sentences over and over again. The text maintains the quality of corpus, reducing the re-source frequency swelling.
Atomic Write (Atomic Write)
The technique to write a data file to a temporary target and then move it to the original target in one move, especially in notes and JSON outputs, to ensure data security.
JSON
Light text format that stores data with domain-name and value pairs. New word definitions are used for new meaning findings and export packages.
Not: This dictionary is a list of living areas (Agent-based areas) that will be added later in the future (Ajan-based fields, Multi-Agencies, Information graphs, etc.). TDK and academic publications for Turkish equivalents of termies have been referred to; widely used in controversial responses.
6. Statistics and Interpretation
Frequency
How many times does one word go through corpus in total? The absolute value should be interpreted with the other word frequency for comparison.
Number of Forms
How many different surface shapes of a word is corpus. Example: gelmek root; Coming, coming, coming I'm coming... can occur in dozens of forms.
Token vs. Lemma
Token: Every word sample in the text.
Lemma: separate dictionary clause.
Example: geldi, geliyor, gelecek separate token, but one lemma: gelmek.
Coverage Rate
The number of dictionary enterry in GTS is at least once passed by corpus. Low rate: the word of the dictionary is less used in archaic or contemporary Turkish.
Token Coverage
How many percent of all the word samples in the collection match GTS matter? High rate: the majority of the text in the compilation consists of standard dictionary words.
Trend Indicators
Rising: word corpus is tightening towards the end (may be new use).
Sabit: Regular use throughout all corpus.
T-BDLD Corpus Statistics
The instant numerical indicators of Turkish context-sensitive Lemmatization Corpus (T-BDLD). Data is automatically updated every time a new analysis is completed.
T-BDLD Summary
Loading statistics...
Lemma Frequency Distribution
Distribution of Inflected Surface Forms per Lemma
Most Frequent Lemmas
Most Frequent Surface Forms
7. Frequently Asked Questions
▸How do you determine the new word candidate?
▸How does the discovery of new meaning work?
New-Meaning Discoveries The results on the page are candidates produced by comparing existing meanings to corpus (T-BDLD) and GTS. The system examines the new contexts, trigger words and co-existences of existing dictionary enterry heads in the collection; The new meaning, meaning expansion, or use candidate for monitoring prepares a structured record.
This process is heavy, so it can be started manually from the management field and stopped; corpus is recommended to be run at non-inclusion hours. Descriptions are prepared as meaning statements, not in the form of a pattern, but in accordance with the modern vocabulary criteria, and presented to expert dictionary control.
Note: The records on this page are automatic candidate inventory. The appearance of a recording does not mean it will be taken directly into the dictionary; the decision is made through the editor's review.
▸Why is the context index important?
▸What's the difference between GTS/YS/D badges on the morphology page?
- GTS (red): word Advanced is available in English Dictionary as dictionary entry.
- YS (green): There is only corpus for the new Word Island, and the quality to be evaluated for the dictionary.
- D (blue): corpus passes, but the candidate has not passed through the filter (special name, attraction, etc.).
▸What happens when corpus is updated? How long do I have to wait?
When the corpus (T-BDLD) is updated, the loading, cleaning and adding phases are monitored in the management interface. The main analysis can be started manually by the manager. Also, if the timed corpus update is open, the system works the same proactive update path as manual analysis by reading the last part added after the final analysis (the default 03.30) at the time set with Turkey clock. Full-to-nothing scan is done only when asked openly. As long as this time mission has its own key, the master automation will kick in even if it's turned off. Progress bar And the instant analysis steps appear: reading corpus, ▪ one-word candidates ▪ multi-word candidate ί context ο morphological families are compiling the results.
In the meantime, users wait for a while The analysis period depends on the size of corpus. Search and other pages continue to work with previous results during analysis; only new word/sympathetic lists and derivatives remain old until families are updated.
When analysis is complete, all pages are available. automatically displays new results; You just have to renovate the page.
▸I saw the update bar, what should I do?
▸What Is the Phrase-Formation Ratio, and What Does the Purple Link Beside a Word Do?
Conceivability rate, for example, shows how much of the total frequency of the single word is in multiple-word clusters. "Uncomfortable" If the word goes by 1,000 times and 250 times "indisposed" or "inappropriate." If you see it in clusters, you'll see it at 25%.
If the word shows high massing, it's on the list on the home page. It's purple and highlighted. (multiword number) and "Absorption rate: %N" rozetleri bulunur.
Clicking on the word is below The panel that opens down It appears that the entire multiword expression is listed as the chip. If you click on Chip, that multiword expression will be placed in the search box and the filter will be applied to the list. If click on the word again, the panel will close.
▸What Is the Difference Between a Derivational and an Inflectional Suffix?
- Production supplement: word a new word is derived. Change the word's class (ad/cent/fiil). Example: eye (ad) eyes (fat) glasses (name)Every word derived from the dictionary is the head of dictionary enterry.
- Gravity.: Your word grammar role It determines (person, time, situation, property). It does not change the word's class and create a new word. Example: Glasses, glasses, glassesIn the dictionary, they're all under the same dictionary entry.
System-label corpus auto-classifies both groups for the detailed attachment list. To the concepts section. Look.
▸What Is Verb Voice (Causative, Passive, Reflexive, or Reciprocal)?
There are four verb roofs in Turkish, and they all come with special additions to the verb root:
- Ettirgen (--///-Don't let anyone else do the act. Read it, read it, do it ..
- Edilgen (- I/-il): The act is done by someone else. Do it, see it.
- Returned (--in/in/-n): The subject of the action is directed at him. Wash up, wash up, put on ▪
- Work (- Work/work/ush/shh): The act is mutual. See sight, beat fight
The system detects these attachments with labeled corpus and saves them with words, especially used to monitor new verb derivatives and active roof productivity.
▸What's the page "Page Guide" panel for?
On every page of the system " What Is This Page? " you can open a panel. Click and open a brief summary of the page, where you will see specific term descriptions, use tips and examples.
The purpose of this panel is: Offering an immediate, accessible aid to the page. The general manual (this page) covers the entire system; the page manuals only describe that page. It is particularly useful for new users.
▸What Is the Inflectional Suffixes Tab on the Morphology Page?
Gravity Addendums tabs, all of the filming attachments in the labeled collection (connected controlled) 8 kategoride toplar:
- Ad Durum Ekleri: enjoining, completing, turning, finding, dating, means, equality (-ce)
- Requisitions: 1/2/3 singular and plural ownership
- Contact Additives: 1/2/3 singular and plural
- Plural Equipment: -lar / -ler
- Zaman EkleriThe past is the past, the present is wide, and the future.
- Kip Ekleri: condition, command, demand, necessity.
- Fiilimsiler: participles (-an/-dık/-acak), verbal nouns (-mak/-ma/-ış), converbs (-arak/-ıp)
- Other: negativity (-me), question (-mi), additional action (-) interest --
For every category: Total use, different word count, subtag distribution (loading, interest etc.), most frequent additional formats and sample word The page contains visual distribution (pasta or bar graph); when a single category is selected, the substructures of the category are shown as bars.
Filtreler: Source (Club / Lonely GTS / Lonely New Word / Lonely Corpus) Kategori (all or one category) Search box The name or form of the service works together.
Data source: etiketli_birikim_duzeltilmis_baglam_kontrollu.txt When loaded with corpus, it is effective instantly.
▸What Does the Inflectional-Suffix Breakdown Show in Morphological Neighborhood?
Morphology page ▪ Morphological Neighborhood When you look for a word in the tab, under the parent/brother/tour card Adds to the collection of the word "{word}" panel is added automatically. This panel:
- Searched for Lemma's corpus 8 kategoride gruplar.
- For each label (load, interest, present_time etc.) total number of use and most frequent additional formats (mortal badges).
- Total use of the word is summarized at the top.
Why? To see a lemma's morphological behavior (which are associated with the additionals, which are open to the category of attractions) at a single glance; to measure the gravitational wealth of the word; to identify the very meaningful word different patterns of use.
▸What Do the Source Filters (All / GTS / New Words / Corpus Only) Distinguish?
Three tabs on the morphology page (Family Search, Morphological Neighborhood, Pulls) with the same logic Source filter There is:
- All: corpus, including all the lemmmas passing through.
- Only GTS GTS: Lemmas found in the GTS dictionary as dictionary enterry.
- Lonely New Word YS: Lemmas marked by the system as the new word candidate (not on GTS, no past the filters of the system).
- Only corpus D: corpus passed by, but not yet classified Lemmas are not in GTS, nor is it marked as a new word candidate (special name, acronym, noise, or may not have entered the candidate pool yet).
These filters narrow down word existence morphological behavior within each group It allows you to study separately, for example, which supplements you want to see for the candidates of the YS.
▸What does the cumulative prevalence graph of proverbs, phrases, and combined verbs indicate?
This graph collects direct captures of the patterns in corpus (T-BDLD) of the page or search that appeared at the time in the relevant tab. Shows how many matches there are for the recording, the total number of occurrences, the rate per million words, the distribution of the prevalence band and several of the most frequent recordings.
The graph is not the sum of the entire category archive; it is calculated according to the active list. It can provide additional evidence from the dictionary enter-the-line detailed scan in the very long or in the words that can enter the word.
▸What's the "corpus" tab on the General Search page?
Corpus Tab in the T-BDLD collection of the word you are looking for lemma or (PHP 3, PHP 4) all of his records without context Lists. Every result contains:
- Lemma (word format) total frequency (corpus passed many times).
- (PHP 3, PHP 4): Lemma's surface shapes in the collection, each with its own frequency (or. kalem (589), kalemi (45), kalemler (21)).
- Part of speech badge (ad/fiil/sylf/zarf).
- GTS (red) or Yeni Is the green badge in the word dictionary?
- "impressive." (amber) The badge you are looking for is not the word Lemma but the attractive form (or. kalemler) is a lemma (kalem) is listed below.
Why? The context search is already done in matching dictionary entry/ Meanings tabs (GTS). Text corpus tab, compilation of a word Her morphological behavior (which often used as an inflected form) is used to see quickly. For the dictionary scientist, it is very useful for mapping the existence of words and studying surface format distribution.