Rendered at 03:18:07 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
laichzeit0 8 hours ago [-]
Interesting project. I would love if you could switch fonts to something like New Athena Unicode.
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
spudlyo 6 hours ago [-]
> LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software.
Definitely! I myself have been having a lot of fun[0] recently using LLMs to enrich one of my favorite beginner Latin readers[1] with audio forced-alignment, POS tagging, morphology, definitions, and other niceties that learners might appreciate.
Kinda font-related, it's throwing me off how every V in the Latin texts show as U. I don't like how they modernized the spellings in school, but if you're not doing that, aren't they all V?
Also most of the word popups use different spellings, like "iam" becomes "jam".
retrac 3 hours ago [-]
Using either spaces or miniscule letters, is anachronistic for both classical Latin and Greek. They only had a single case in antiquity and didn't usually mark word boundaries.
Sort of. Romans didn't handwrite like that. See https://en.wikipedia.org/wiki/Roman_cursive . Note that the "u" character in both ancient and late ancient Roman cursive is more like the modern "u".
(Source: not an expert in any way)
usern20260720 4 hours ago [-]
My favourite ancient greek font is Theano
tmshapland 5 hours ago [-]
I'm always surprised to see that a large enough portion of the HN community are interested in classics that posts like this make it to the front page. Who are you all? What are you doing here?
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
How did you get here?
Curious_Cake 4 minutes ago [-]
I read quite a lot as a kid. Fiction and non-fiction. I'm intensely curious, and the "why" and the "how" of things leads you into new subjects. It doesn't hurt that history is also filled with stories as compelling as most fiction. Much of the old stuff also holds up quite well. Meditations by Aurelius is amazing. Also, Confessions by Augustine of Hippo.
I have many interests but botany became a big one, and the scientific naming scheme is mostly Latin and Greek, which leads to some interesting word history. "Sativa" means anything "cultivated" (example, "Avena sativa" is "cultivated oats" and "Crocus sativa" is "cultivated saffron"). "Nemo" means "no one" (so the "Search for Nemo" is the search for no one, haha), the root word for "service" means "slave". There are many more interesting associations.
I find the deepest and most awe-inspiring truths within physics (out of which biology and chemistry are arguably sort of abstractions in terms of order of complexity). Behind this is a long awkward history of curious people fumbling for answers in the dark. I am an odd nut, but it is connecting in a comforting way the way I can see others have asked the same burning questions of the "how" and the "why" of things, and the myriad subjects it has lead them to.
As John Muir once said: "When we try to pick out anything by itself, we find it hitched to everything else in the universe."
ocd 47 minutes ago [-]
When you reach an age in your 20s or maybe 30s where you may begin to actually understand the concept of time, these subjects become like catnip.
spudlyo 4 hours ago [-]
I'm a recently retired software nerd who got into reading 19th century literature. I was greatly inspired by characters in these novels, most of whom had the advantages of a classical education, and I lamented how poor my own was. One day I idly wondered how hard it would be to learn a bit of Latin (as I was constantly having to look up Latin references) and once I discovered LLPSI[0] I was hooked.
I grew up with an awareness of Greek and Latin as part of western heritage education.
A lot of role playing games, especially from Square Enix,leverage a lot of Latin and classical education.
yardshop 3 hours ago [-]
I became more interested in Latin when learning to sing the Carmina Burana, after already knowing some German and Spanish. I haven't learned much beyond that, but still fascinated by it as with all languages.
bonoboTP 4 hours ago [-]
I'd think there's a strong correlation. It's a kind of nerd stuff. Greek mythology, classics, train schedules and types, all the space probes and all the NASA stuff, Haskell etc etc. It's all very similar.
2 hours ago [-]
sfRattan 5 hours ago [-]
Latin in school from 6th grade onward. And a classical education with many holes and gaps that I am continually filling in.
ButlerianJihad 5 hours ago [-]
Practicing Roman Catholic family with liberal arts education
nephihaha 5 hours ago [-]
I think the answer is that the past is as relevant as the future. If people here know both computer programming and classical history, I think that means they have a more rounded view of the world. (The problem in my part of the world, unfortunately is that classical education has often been class based.)
The thing that strikes me from ancient texts is not the obvious differences from our day, but the similarities.
May I point out that in "Crudam si edes, in acetum intinguito", that "edes" is more likely to be the future of edere/esse "to eat"? (Just guessing by context.)
lsb 4 hours ago [-]
Fixed!
svat 7 hours ago [-]
I remember encountering NoDictionaries years ago, and love it. Thank you for building it!
beloch 7 hours ago [-]
Suggestion: For the pop-up text, bold the meaning of the word so it pops out a bit better. Due to the formatting of word definitions, you have to go hunting for it in many cases.
ndoddhdudhhd 2 hours ago [-]
And make it faster to load!
If these 2 are fixed I’d use it daily
frollogaston 6 hours ago [-]
Also sometimes you have to expand the entry to get to the meaning at all
b3orn 6 hours ago [-]
And it's not immediately obvious how to close a pop-up again.
smithkl42 6 hours ago [-]
Love this.
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
I recognise that quote. :) But am I right in noticing a lack of breathings? Unless the tilde indicates it. I'm not sure of the difference between the acutes and graves here. I am used to seeing only one kind of accent on ancient Greek vowels.
DonaldFisk 4 hours ago [-]
The breathing marks are clear in the Ancient Library texts (e.g. St John's Gospel), but difficult to make out on Hacker News.
A propos of breathing marks, iota subscripts, and the three different accent marks of Classical Greek, when I learned it at school we had to remember the breathing marks and iota subscripts, and would lose marks if we omitted them, but we didn't need to learn the accents. Modern Greek now has only (acute) accents, which you need to know to stress the correct vowels, exactly where the accents were in the equivalent Classical Greek words.
mkehrt 1 hours ago [-]
There are in fact rough breathing marks there on the articles, stacked under the inverted breve accents. Which is an alternate form of tilde.
milkcrate 5 hours ago [-]
Could vary by text? I see breathing marks in Homer.
retrac 2 hours ago [-]
Modern editions usually use modernized spellings. There are no breathing marks in the oldest manuscripts; they are a classical innovation (circa 400 BC) several centuries after Homer was first written down (600 BC or earlier). Though most of Homer as we receive it today is through medieval manuscript copies, and copyists often "corrected" the spellings just like we do today.
soiltype 8 hours ago [-]
Does this have some use case that the Perseus Digital Library doesn't already serve?
flats 7 hours ago [-]
I was wondering the same thing —Perseus has been around since _1987_ & seems to have more features? I guess this is a bit more attractively laid out…
thaumasiotes 6 hours ago [-]
Note that Perseus considers itself end-of-life; there is an analogous project that's supposed to have succeeded it, but I don't remember what project that is.
milkcrate 5 hours ago [-]
It's called Scaife[1] and it's borderline unusable. Slows to an absolute crawl on my ~2020 machine within a few minutes, and the viewing experience is very uncomfortable. I'd be really surprised if people preferred it to Perseus.
Yeah, it's awful. I can't believe how much of a downgrade it is from the old Perseus Hopper.
equalbeforegod 4 hours ago [-]
Agreed. Scaife supporters are delusional.
What should actually be done, and what OP should do, is take the Perseus website and make it so it doesn't 503 all the time (it is incredibly unreliable).
Then, FOIA the State of California for the pay-to-play data which the UC Irvine-based TLG hoarders are withholding from the public (it is a publicly funded project...) so that the corpus can be meaningfully extended and built upon.
The 1990's tier html vibe of Perseus is to its great advantage.
6 hours ago [-]
cwnyth 6 hours ago [-]
Is it not just Perseus with a wrapper?
Transformanshen 3 hours ago [-]
An interesting project. Ancient languages have always fascinated me, but the effort required to read even a single page of Greek or Latin can be a serious obstacle
It's nice to see tools that simplify this process. I might even want to give it another try
turpentine 3 hours ago [-]
Online tools to simplify this process like Perseus have existed for at least 15 years, and you could view original texts side by side with English for comparison, and original words have all their cross-references hyperlinked etc. All maintained by academics committed to rigorous and nuanced approaches to translation.
Ancient Library is some vibecoded alternative that the LLM translation it uses will be incorrect. Color me unimpressed.
usern20260720 4 hours ago [-]
I generated a read-along version of Athenaze read by a real native language Greek speaker and I met some of the problems that this web has: the vast amount of vocabulary makes managing a dictionary rather complicated.
For tbos version, I gather that a bilingual presentation would be more than necessary: keep Ancient Greek text on a side and display a scholar translation into a selection of switcheable languages
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
Planktonne 7 hours ago [-]
William Whitaker's Words [1] is the best resource I've ever come across for this. 'Fulgere' is 'to shine' [2].
That's actually ok. In virtually all Latin dictionaries verbs are listed in the first singular person of the present indicative (e.g. you would find "sum" and not "esse"). It's just showing you the dictionary entry.
marginalia_nu 7 hours ago [-]
It's not really very useful in this particular case, as you can't actually navigate the "dictionary" any other way than finding the base word form somewhere in the text you're looking at.
6r17 6 hours ago [-]
I built myself such a simple lexicon for technical stuff (concurrency vocab - invariant, genetics, stuff like that) - can only recom the practice as vocabulary is clearly a big step difference
6 hours ago [-]
veqq 6 hours ago [-]
Vowel lengths on the main text would be great.
leoc 1 hours ago [-]
I'm a bit out of energy at present but I'll try to put together a comment.
From the point of view of a learner it's unfortunate that macrons (diacritical marks for long vowels) have not been added to the text, and that 'u's have not been altered to 'v's where appropriate.
The word-by-word treatment of translation makes this broadly a part of the category of interlinear texts. For comparison here's one of the early pages of Max Müller's 1864 Sanskrit-to-English interlinear of the https://en.wikipedia.org/wiki/Hitopadesha : https://archive.org/details/firstbookofhitop00ml/page/2/mode... . This is part of a revival of interlinear texts, as a resource for language learners, which traces back to John Locke but really caught on in the English-speaking world in the early nineteenth century thanks to James Hamilton. It's unusual for "traditional" interlinears in having one row which gives a(n explicit) morphological gloss of each word in the original (the "da, 3 sg. Pres. Par" and so on) in the same column as the original word, an English translation of the word, and (in this case) a transliteration. But morphological glosses are a standard feature of modern "interlinear morphemic glosses" https://www.christianlehmann.eu/ling/ling_meth/ling_descript... made by and for linguists who want to analyse and compare languages, rather than learn them. Linguists' adoption of modern IMGs seems to have taken off in the 1960s (though some use has been made of interlinears for "serious" linguistic purposes since at least the 1890s beginnings of the Linguistic Survey of India https://en.wikipedia.org/wiki/Linguistic_Survey_of_India ).
it would be nice to sort by date or date range as well
hcayless 4 hours ago [-]
The Greek font is terrible—it’s using the Modern Greek “acute” accent which looks weird, and commas aren’t supposed to have spaced before them. It’s cool that Perseus did the work that makes this possible though. More info about what edition we’re looking at would be nice too.
gaigalas 6 hours ago [-]
Clicking out to close a dictionary popup only works if you click empty space within the layout center. This is annoying. On a wide monitor, most empty space is on the sides (where clicking doesn't close).
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.
gfaure 7 hours ago [-]
Cool! Enclitic -que should _not_ be broken off as a separate word, though.
cwnyth 5 hours ago [-]
That plus the spaces between punctuation makes it clear it's done for tokenization, but nothing puts it back together again.
frollogaston 6 hours ago [-]
Wow, I used to spend forever looking up words in Latin books.
Boss0565 7 hours ago [-]
I like it. Can you make the formatting on parsed words easier to read?
5 hours ago [-]
milkcrate 5 hours ago [-]
Could stand to have a better Greek font, but this is great stuff. Thanks for sharing.
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
Definitely! I myself have been having a lot of fun[0] recently using LLMs to enrich one of my favorite beginner Latin readers[1] with audio forced-alignment, POS tagging, morphology, definitions, and other niceties that learners might appreciate.
[0]: https://hercules.hookbangsplat.com
[1]: https://archive.org/details/p1fablesoforbili00godl/mode/2up
Also most of the word popups use different spellings, like "iam" becomes "jam".
For example:
https://en.wikipedia.org/wiki/File:Trajan_inscription_duoton... (Latin as typically carved in stone)
https://commons.wikimedia.org/wiki/File:Herculanean_Rolls_-_... (Greek as typically written on papyrus)
(Source: not an expert in any way)
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
How did you get here?
I have many interests but botany became a big one, and the scientific naming scheme is mostly Latin and Greek, which leads to some interesting word history. "Sativa" means anything "cultivated" (example, "Avena sativa" is "cultivated oats" and "Crocus sativa" is "cultivated saffron"). "Nemo" means "no one" (so the "Search for Nemo" is the search for no one, haha), the root word for "service" means "slave". There are many more interesting associations.
I find the deepest and most awe-inspiring truths within physics (out of which biology and chemistry are arguably sort of abstractions in terms of order of complexity). Behind this is a long awkward history of curious people fumbling for answers in the dark. I am an odd nut, but it is connecting in a comforting way the way I can see others have asked the same burning questions of the "how" and the "why" of things, and the myriad subjects it has lead them to.
As John Muir once said: "When we try to pick out anything by itself, we find it hitched to everything else in the universe."
[0]: https://en.wikipedia.org/wiki/Lingua_Latina_per_se_illustrat...
A lot of role playing games, especially from Square Enix,leverage a lot of Latin and classical education.
The thing that strikes me from ancient texts is not the obvious differences from our day, but the similarities.
May I point out that in "Crudam si edes, in acetum intinguito", that "edes" is more likely to be the future of edere/esse "to eat"? (Just guessing by context.)
If these 2 are fixed I’d use it daily
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
ΕΝ ΑΡΧΗ ἦν ὁ λόγος , καὶ ὁ λόγος ἦν πρὸς τὸν θεόν , καὶ θεὸς ἦν ὁ λόγος .
A propos of breathing marks, iota subscripts, and the three different accent marks of Classical Greek, when I learned it at school we had to remember the breathing marks and iota subscripts, and would lose marks if we omitted them, but we didn't need to learn the accents. Modern Greek now has only (acute) accents, which you need to know to stress the correct vowels, exactly where the accents were in the equivalent Classical Greek words.
[1]: https://scaife.perseus.org/
What should actually be done, and what OP should do, is take the Perseus website and make it so it doesn't 503 all the time (it is incredibly unreliable).
Then, FOIA the State of California for the pay-to-play data which the UC Irvine-based TLG hoarders are withholding from the public (it is a publicly funded project...) so that the corpus can be meaningfully extended and built upon.
The 1990's tier html vibe of Perseus is to its great advantage.
Ancient Library is some vibecoded alternative that the LLM translation it uses will be incorrect. Color me unimpressed.
For tbos version, I gather that a bilingual presentation would be more than necessary: keep Ancient Greek text on a side and display a scholar translation into a selection of switcheable languages
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
[1] https://latin-words.com
[2] https://latin-words.com/word/latin/fulgere
From the point of view of a learner it's unfortunate that macrons (diacritical marks for long vowels) have not been added to the text, and that 'u's have not been altered to 'v's where appropriate.
The word-by-word treatment of translation makes this broadly a part of the category of interlinear texts. For comparison here's one of the early pages of Max Müller's 1864 Sanskrit-to-English interlinear of the https://en.wikipedia.org/wiki/Hitopadesha : https://archive.org/details/firstbookofhitop00ml/page/2/mode... . This is part of a revival of interlinear texts, as a resource for language learners, which traces back to John Locke but really caught on in the English-speaking world in the early nineteenth century thanks to James Hamilton. It's unusual for "traditional" interlinears in having one row which gives a(n explicit) morphological gloss of each word in the original (the "da, 3 sg. Pres. Par" and so on) in the same column as the original word, an English translation of the word, and (in this case) a transliteration. But morphological glosses are a standard feature of modern "interlinear morphemic glosses" https://www.christianlehmann.eu/ling/ling_meth/ling_descript... made by and for linguists who want to analyse and compare languages, rather than learn them. Linguists' adoption of modern IMGs seems to have taken off in the 1960s (though some use has been made of interlinears for "serious" linguistic purposes since at least the 1890s beginnings of the Linguistic Survey of India https://en.wikipedia.org/wiki/Linguistic_Survey_of_India ).
https://theamericanscholar.org/the-new-old-way-of-learning-l... is a very incomplete history of interlinears (the French are especially shortchanged) but probably still the best thing out there in English. See also https://www.reddit.com/r/interlinear . (I have to release some things about the history of interlinears myself but I'm years late at this point.)
https://shop.hyplern.com (might as well add an affiliate link! https://invi.tt/N5VZXP7g ) and https://interlinearbooks.com/ are two publishers offering modern interlinears aimed at language learners. The polished but expensive Legentibus service https://legentibus.com/ offers some interlinears too.
Guess who Vincent F. Hopper https://archive.org/details/chaucerscanterbu0000chau_k0e4 was for a time married to!
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.