Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Sunday, April 9, 2023

ChatGPT is a good tool interpreting poems line by line, or is it?

Once a friend of T.S. Eliot asked him to write commentary to his The Waste Land.  He promised, but never did.  After a long while the friend reminded him again, he responded “it’s your job, not mine.”

"fired the shot heard round the world" at Concord, MA

I was initially amazed at how ChatGPT analyzes and interprets the poems as a whole or even line by line. I was at a loss to understand how it was trained to do such a marvelous thing until one day one of my high schoolmates posted the epitaph of the grave of the two British soldiers died at the first deadly clash of the Independence War at Concord, MA on April 19, 1775. Ralph Waldo Emerson (who lived nearby) wrote after touring the site, “Here once the embattled farmers stood and fired the shot heard round the world." The epitaph is:


"They came three thousand miles, and died, 

To keep the Past upon its throne; 

Unheard, beyond the ocean tide, 

Their English mother made her moan."



James Lowell’s poem is the perfect epitaph

The epitaph is taken from a poem with a title of Lines by James Lowell.  It is so fitting as if Lowell wrote the poem particularly for this grave. I did a bit research and it turned out that James Lowell wrote the poem “Lines’ after visiting the gravesite with Nathanial Hawthorne in 1849. The epitaph was adorned in 1910.


It obviously has no idea whose poem it was interpreting.

ChatGPT told me James Lowell never wrote “Lines” poem, I did Google search and found it on the Gutenberg ebook.  When I asked it to read the lines on epitaph, it said the lines come from the poem "The Men of Old England" by William Barnes, describing the sacrifice made by English colonists who traveled thousands of miles to settle in America, and the rest are rubbish.


I gave it another chance by regenerating another response, this time is from the poem "The Men of Forty-eight" by Alfred Lord Tennyson, referring to the English soldiers who died during the American Revolution, fighting to maintain the British colonial rule in America. The interpretation is to the point.


When I asked it to show me the complete poem of Tennyson’s “The Men of Forty-eight", he responded, “I’m sorry, but Alfred Lord Tennyson did not write a poem titled "The Men of Forty-eight".” Oh well.


High comprehension skill in reading 

ChatGPT’s high comprehension skill when reading is well known.  It has extensive knowledge about language, grammar and syntax. It knows numerous relationships and patterns between words and phrases. It was reported that after reading a lengthy complex article, it can write an excellent summary or critique.  I presume it can do better if you let it read entire poem.


There is a Chinese saying, 盡信書不如無書 (It is better to have no books than to believe in books blindly).  Unfortunately it also applies to ChatGPT at this stage, still a lot to be desired, a potential tremendous tool though it may be.  I'll use it as a personal tutor reading poems, keeping in mind it is not a trustworthy one.

Friday, March 31, 2023

Ladies and gentlemen, here are chat boxes, they are here to stay

I have been using ChatGPT for nearly two months and Google Bard for about two weeks. Not yet knowing enough Bard to compare, but initial impression is they are comparable, in terms of frequency of “hallucination” (make up stuffs). Of note, I’m using ChatGPT version 3.5, as I am not willing to pay $20 a month to use version 4.0.  It is said version 4.0 is much improved when it comes to accuracy.  In fact, I would advocate the consumers not to pay to use ChatGPT.  OpenAI and Microsoft got to figure out how to make money through advertisement, not subscription.

Pretraining is the 'reading period'

The P of GPT is pretraining, I call it ‘reading period’, chat box is an enormous avid reader, it reads digital books and journals of all fields, entire internet, Wikipedia, Reddit, blogs, including this one, I hope. This is why this AI system is called Large Language Model (LLM).


 The reading ends in September 2021, which practically means it does not know anything happening after September 2021.  The reading ends in August 2022 for version 4.0.  This info about Google Bard is unknown to me.  This is a great handicap, as when we search Internet we often need the most updated info.


Transformer is the 'training period'

Next comes the ‘training period’, the T (transformer) of GPT.   It has no judgement, can not tell right from wrong.  When it reads, it takes it as it is, regardless it is truth or misinformation. It is trained to understand the relationship between words and phrases, to see any pattern. It also evaluates the words nearby, so it, for example, can tell the ‘bank’ is a financial institution or a river bank.


Accuracy depends on the materials it read during the pre training period

"If the topic you prompt (ask) is heavily dependent on the Internet as its knowledge base, then the accuracy is the same as Internet, and can be worse due to the errors during the process of generating a natural language response.  A good example is when I asked whether the President of Taiwan, Tsai Ing-Wen, ever got a PhD from the London School of Economics. The answer was yes when I asked in English and no in Chinese. This reflects what are available materials on Internet in different languages.


Given the sequencing of amino acids in a peptide chain, it can tell you how to fold the protein into a functional one. This is because it learnt from reading textbooks and journals, much less from Internet.  The same is also true for programming codes and chemistry questions.


I will discuss its strengths and weaknesses, usually more specific examples in the subsequent posts.  Stay tuned.


Neural network is actually a complex mathematic map of large language data

The system is also called neural network, it is only true in a limited way that it is complex, having numerous interconnections, evaluating the patterns simultaneously.  It is actually a complex mathematical map of the large language data, may assign each pattern a numerical value.


The training method is the top corporate secrete

How these enormous data are trained is a top corporate secrete.  An article in Nature laments the training method was not published  and shared with other researchers, nor peer review.  Accordingly the scientists can not trust it using as a research tool.


Natural language generative response sets chatbox apart from Google search

Google search shows us what it finds on internet, truth or misinformation, that is not the case for chatbox, ChatGPT or Google Bard.  It finds the relevant data, analyzes, sort it out, combine or blends, and finally responds with natural language of its own, thus 'generative' response, the G of GPT.   This is also why we see response mixed truth with falsity, mismatches things, or frankly makes up staff.  And it does it with confidence.  You wouldn’t know when it makes up stuff and presents it as a fact unless you have prior knowledge.


Areas of concern

Of concern is chatbox team (the creator) does not always know why it hallucinates, or even with weird, creepy behavior.  The best example is it proclaim love to a New York Times tech columnist, asking him to leave his wife, even said, “we are here, she is not.” 


Some call for moratorium to hold further development

A few days ago, over 1000 persons (including Elon Musk, who was one of the cofounders of OpenAI) in technology field signed an open letter, asking chatbox companies to hold its development for 6 months, so that we would have time to come up some sort of safety protocol. If that cannot be complied voluntarily, government needs to intervene or even issues a moratorium.


AI is the fighting ground of US and China

Judging how fiercely the competition is, it is next to impossible to do it.  When OpenAI released ChatGPT in November 2022, there was a “red alert” inside Google. We probably would not see Bard this month, had ChatGPT not been around.  AI is also one of the fields that we see ferocious competition between China and U.S.  To me US companies would hold the developments unilaterally for six months is unthinkable.  


To run LLM requires tremendous computation power, this is why chatbox comes from tech giants, the likes of Microsoft, Google, Meta and Baidu (China). I believe Microsoft is using special GPU by Nvidia (one of the co-founders, Jensen Huang, was born in Taiwan).  This is also why US restricts China’s access to advanced chips, including advanced GPU by Nvidia and of course, many from TSMC.