Wednesday, February 19, 2014

Blog Post #3- Exploring Text Visualization Tools


People kept dying in February!
One of them was artist Nancy Holt, and after reading about her life/death in the Times, it got me wondering about the language of obituaries. So for my blog post #3, I compared a dozen women’s and men’s obituaries in the NY Times, using Wordle and Voyant- which was much more interesting!
First the issues- it’s hard to compare people- it’s subjective, arbitrary- so I tried to find pairs that would be the most like comparing apples to apples. 
Like: same profession, similar length of life, comparable volume of output/influence, similar kinds of work, similar institutional/popular recognition, generally recognized on an equal level and similar length of obituary. I also tried to get relatively recent obituaries- except for Georgia O’Keeffe (1986) and Clement Greenberg (1994)- most (the other 10) are from the past decade.
I KNOW, that’s pretty shaky ground to start on-but I was having fun with the idea and wanted to use the tools to see what I could see in the texts.
First I did them one at a time- one man, one woman- and these were the pairing categories:
Earthwork/Land Art artists: Nancy Holt and Walter De Maria
Big famous artists whose work appears on mugs and totebags:  Georgia O’Keefe and Robert Rauschenberg
Big famous artists who aren’t average kitchen table names: Louise Bourgeois and Cy Twombly
Big famous people you read in art theory class: Susan Sontag and Stuart Hall
Big famous critics: Ada Huxtable and Clement Greenberg,
Big famous Nobel Prize winning writers: Doris Lessing and Seamus Heaney. (Here I wanted to do poet vs poet –(Wislawa Szymborska) but thought maybe it’s better to compare English language literature since it’s an English Language newspaper??)
These are the results from a few individual comparisons in Wordle: 
Nancy Holt and Walter De Maria:




Doris Lessing and Seamus Heaney: 

Susan Sontag and Stuart Hall:


Then I did them all together:
Women: Doris Lessing, Ada Louise Huxtable, Nancy Holt, Louise Bourgeois, Georgia O’Keefe, Susan Sontag


Men: Seamus Heaney, Clement Greenberg, Walter De Maria,  Cy Twombly,  Robert Rauschenberg, Stuart Hall


I also put them in Voyant tools because I wanted to compare how many action words/verbs were in the pairings, but found Voyant hard to use in that respect. I was able to get word counts for each word, but found it hard to select a grouping of words (like “made, wrote, said” all together) because the texts were long and the panel called Frequencies Word Corpus seems to only allow one page at a time (?)
So I tried it a little bit differently. I put in the words “work” and “works” for a few one on one comparisons. Here are the results.


The word work(s) in obituaries of Nancy Holt and Walter DeMaria- her obit is shorter, but more mentions of work/works- but if you count earthworks, he has a few more.

The word work(s) in obituaries of Doris Lessing and Seamus Heaney- her obit is longer, but with fewer mentions of work.
Then, there was this- a big difference in occurance of the word work(s) in obituaries of Louise Bourgeois and Cy Twombly: 
 
Next, I put all the Women’s obituaries together and all the Men’s obituaries together, creating a larger corpus, and did the same search- there was a noticeable difference, even though the total word counts were relatively similar. 
Women Total word count: 10,228; work(s) mentioned:34 
Men Total word count: 10,405; work(s) mentioned: 57


I think what I learned here is that the corpus has to be big to really see patterns evolve. It would be great to do this on a really large scale, to see if the same patterns would get larger. (Reinforces the Moretti article.)
In terms of students doing this: This was fun to do, there’s a visual pleasure right away with Wordle, but to get a little deeper, with Voyant, took time and effort and also a deep curiosity on my part. I think you have dedicate time to learn how to use the tools and start out with a good question- a question that leads to more questions. I think students would have to use material they are deeply interested in to start the process of a committed exploration- and then see all the possibilities of looking at the data from different angles.
PS: I also tried Docuburst, but I’m not exactly sure how to read these results accurately. 

Monday, February 3, 2014

Data Visualization is Fun?

Somewhat related to the Moretti article I read, here is a link to a fun sight of data visualizations gone wrong.


Blog Post #2- Challenges and Opportunites in the DH


The Moretti article presents the idea of using graphs to trace the rise of the novel in a few different cultures (mostly European) and discusses how the graph data has been used correctly or incorrectly for various interpretations. He discusses how literary criticism- which has not really had a use for quantitative analytics- could be enhanced and effected by incorporating this type of data. He starts out by saying the traditional models of literary analysis (comparing it to older methods of historical research) focused on exceptional events as the primary area of focus, but Moretti is suggesting that is all the other stuff- the not so extraordinary, the everyday and mass data- can be even more useful. In this he is suggesting an opportunity for better, deeper understanding.
He describes the benefit of big data/graphs: “… a field this large cannot be understood by stitching together separate bits of knowledge about individual cases, because it isn’t a sum of individual cases: it’s a collective system, that should be  grasped as such, as a whole—and the graphs that follow are one way to begin doing this.”
What he is suggesting is that there is a greater value in looking at larger data sets, rather than isolated incidents-more macro than micro- because the larger data can reveal patterns and systems which are likely to be more useful and accurate in research.
He also (importantly) mentions that the gathering of such data is a collective and shared effort, and that the scope of this kind of research benefits from collaborations since results can be combined in multiple ways.
He supports his ideas by examining the how rise and falls of genre novels can be overlayed on the more general “rise of the novel” measures, and also how other factors (politics, generational shifts, trends) can affect the preliminary assumptions.
Because the new scholarship he is discussing are graphs, I don’t see him as blowing up any ivory towers. I can’t see any traditional humanities scholar being horrified by this- it’s really just looking at the same information (the rise of the novel) with fresh eyes (quantitative graphs)- taking into consideration other relevant information. I would hope that a traditional humanist would see it as an enhancement rather than a threat to traditional scholarship.
I think students would be very receptive to multiple ways of analysis- especially the ones with visualizations, but I think it’s important to have them stand in relationship to other sources. Data is most useful when viewed through multiple contexts. For example: English books imported in India as a graph relates one thing. Considering other data like information of the political climate, the cost of paper at the time, even sailing data patterns etc., changes the way we would read the data.
A graph, although we like to think of them as neutral or unbiased is so open to interpretation, if it’s isolated as a singular source. As pointed out in the article, It’s the layering of information gathered by the graphs that provides a real understanding.