Wednesday, February 19, 2014

Blog Post #3- Exploring Text Visualization Tools


People kept dying in February!
One of them was artist Nancy Holt, and after reading about her life/death in the Times, it got me wondering about the language of obituaries. So for my blog post #3, I compared a dozen women’s and men’s obituaries in the NY Times, using Wordle and Voyant- which was much more interesting!
First the issues- it’s hard to compare people- it’s subjective, arbitrary- so I tried to find pairs that would be the most like comparing apples to apples. 
Like: same profession, similar length of life, comparable volume of output/influence, similar kinds of work, similar institutional/popular recognition, generally recognized on an equal level and similar length of obituary. I also tried to get relatively recent obituaries- except for Georgia O’Keeffe (1986) and Clement Greenberg (1994)- most (the other 10) are from the past decade.
I KNOW, that’s pretty shaky ground to start on-but I was having fun with the idea and wanted to use the tools to see what I could see in the texts.
First I did them one at a time- one man, one woman- and these were the pairing categories:
Earthwork/Land Art artists: Nancy Holt and Walter De Maria
Big famous artists whose work appears on mugs and totebags:  Georgia O’Keefe and Robert Rauschenberg
Big famous artists who aren’t average kitchen table names: Louise Bourgeois and Cy Twombly
Big famous people you read in art theory class: Susan Sontag and Stuart Hall
Big famous critics: Ada Huxtable and Clement Greenberg,
Big famous Nobel Prize winning writers: Doris Lessing and Seamus Heaney. (Here I wanted to do poet vs poet –(Wislawa Szymborska) but thought maybe it’s better to compare English language literature since it’s an English Language newspaper??)
These are the results from a few individual comparisons in Wordle: 
Nancy Holt and Walter De Maria:




Doris Lessing and Seamus Heaney: 

Susan Sontag and Stuart Hall:


Then I did them all together:
Women: Doris Lessing, Ada Louise Huxtable, Nancy Holt, Louise Bourgeois, Georgia O’Keefe, Susan Sontag


Men: Seamus Heaney, Clement Greenberg, Walter De Maria,  Cy Twombly,  Robert Rauschenberg, Stuart Hall


I also put them in Voyant tools because I wanted to compare how many action words/verbs were in the pairings, but found Voyant hard to use in that respect. I was able to get word counts for each word, but found it hard to select a grouping of words (like “made, wrote, said” all together) because the texts were long and the panel called Frequencies Word Corpus seems to only allow one page at a time (?)
So I tried it a little bit differently. I put in the words “work” and “works” for a few one on one comparisons. Here are the results.


The word work(s) in obituaries of Nancy Holt and Walter DeMaria- her obit is shorter, but more mentions of work/works- but if you count earthworks, he has a few more.

The word work(s) in obituaries of Doris Lessing and Seamus Heaney- her obit is longer, but with fewer mentions of work.
Then, there was this- a big difference in occurance of the word work(s) in obituaries of Louise Bourgeois and Cy Twombly: 
 
Next, I put all the Women’s obituaries together and all the Men’s obituaries together, creating a larger corpus, and did the same search- there was a noticeable difference, even though the total word counts were relatively similar. 
Women Total word count: 10,228; work(s) mentioned:34 
Men Total word count: 10,405; work(s) mentioned: 57


I think what I learned here is that the corpus has to be big to really see patterns evolve. It would be great to do this on a really large scale, to see if the same patterns would get larger. (Reinforces the Moretti article.)
In terms of students doing this: This was fun to do, there’s a visual pleasure right away with Wordle, but to get a little deeper, with Voyant, took time and effort and also a deep curiosity on my part. I think you have dedicate time to learn how to use the tools and start out with a good question- a question that leads to more questions. I think students would have to use material they are deeply interested in to start the process of a committed exploration- and then see all the possibilities of looking at the data from different angles.
PS: I also tried Docuburst, but I’m not exactly sure how to read these results accurately. 

10 comments:

  1. Oooohhhhh..... fun idea!!! The difficulty in comparing obituaries of very different people notwithstanding, the concept is really cool. And I was particularly intrigued by the Wordle word clouds created when you compared Stuart Hall and Susan Sontag. As someone who is well versed in cultural studies and cultural critique, I have read much Hall and Sontag. The first thing that jumped out at me was that both word clouds featured the term "culture." Given that both offered cultural critiques and analyses, I found this to be particularly relevant in terms of identifying key themes in their lives and work. I also actually thing each word cloud fairly clearly represents each artist and their own body of work. Stuart Hall's famous "Race: The Floating Signifier" was one of the pivotal readings during my doctoral education that influenced my own worldview. And the term "race" and other related terms (e.g., "black") are clearly evident in the word cloud. Similarly, Sontag's critiques about conflict (e.g., Vietnam, Sarajevo) illuminated aspects of war and conflict that had previously been ignored or, worse, silenced. Interestingly, this aspect of her body of work isn't visible in the word cloud (at least, not what I can see; many of the words are too small for my old, tired eyes), but "criticism," "critique," and like terms are scattered throughout the word cloud. This, I think, does represent her body of work because she was an outspoken critic.

    I'm a visual person, a visual learner, I see patterns in things (like Sudoku... and scheduling classes), so the word clouds are, for me, very interesting. I don't really understand the docuburst, so I haven't played with that. Kudos to you for trying something new. And I'm still intrigued by what a fun idea you had to compare obituaries. Woo!

    ReplyDelete
  2. Great idea! If you were to do this on a very large scale, as you discuss in your post, that would be awesome. Imagine a wordcloud that is made of every obituary over the past year. Compare this years word cloud to one from the 1950s or something. People will flock to your blog for sure to see that.

    ReplyDelete
    Replies
    1. Thanks Glenn- I think the idea of factoring in the year of death is a really great one because it would (maybe) show how the language of obituaries has changed over time too.

      Delete
  3. I found your post very helpful in terms of how to approach this type of research. It's not just a matter of feeding text into a digital tool, but rather your exploration was framed by a series of thoughtful questions, often involving comparisons, that guided you ever deeper into the material. I also wonder if this only works when using very large data sets; as you say, it may take that in order for patterns to be revealed.

    ReplyDelete
    Replies
    1. Yes- the patterns only started to emerge when the data got larger!

      Delete
  4. This seems like such a great idea for an article / classroom exercise. Like Provost Arcario suggests above, it's the questions or framing of this exercise that can get you started. (Interestingly, we *usually* don't know what we'll discover in a visualization. You could obviously 'scale' this up to a larger corpus / corpora with more male and female subjects. The other thing along with what Glenn suggests is that you could try out different sources -- like in Europe, or an artist's home country and see if there are any differences.) But especially for a classroom exercise, this would be really engaging. Take several obits. from a well-known artist / author and see what it was about their life that stood out.... I appreciate that you created a complex corpus and used several different tools. What's great about figuring how to do this once or twice is that it will be so much easier the next time, and obviously, if you want to try these activities in the classroom, you clearly have become something of an 'expert' to assist your students.

    ReplyDelete
    Replies
    1. Yes, the initial results lead to a whole path of possible questions! I have been thinking about it more and more and realizing that once the (large) corpus is established, you can really start to look at it from all these different angles, which then makes you want to refine the corpus in a different way...and so on and so on down the rabbit hole. It was fun though!

      Delete
  5. Great work! Your article seems very interesting. I tried to use large amounts of data in my computer but it didn’t work. I figured out how it should work when I went through your article. Now I could analyze large amounts of data in my math and engineering class.

    ReplyDelete