Показаны сообщения с ярлыком knowledge graphs. Показать все сообщения
Показаны сообщения с ярлыком knowledge graphs. Показать все сообщения

суббота, 8 февраля 2020 г.

5 technology trends for the roaring 20s, part 2: AI, Knowledge Graphs, infinity and beyond

You don't have to be a fortune teller to identify AI as the key trend for the 2020s. But there is nuance regarding AI hardware and software that deserves to be highlighted.

By for Big on Data



Picking up from where we left off with part one of the top technology trends for the 2020s, here is what will shape the data landscape for the years to come.

2. AI: It's all about Data and Hardware

The last part of the 2010s has been all about AI, and the 2020s will not be any different. We will see AI widening its reach, and impacting every conceivable field. Having already seen the AI hype rise, however, we must also be prepared for a backlash. And it's very important to be aware of what "AI" actually means.
In essence, what we call AI today is an umbrella term for various pattern matching techniques. Machine learning and its various subdomains, such as deep learning, essentially boil down to pattern matching. We've seen several breakthroughs in the 2010s, but the seeds for most techniques and algorithms have been planted decades ago and remain essentially the same. 
Still, we have seen the performance of AI systems in many domains going from being worse than human, to catching up and surpassing humans. How is that possible? The answer is twofold: Data and compute.
The digitization of nearly all aspects of human activity has led to an explosion in the volumes of data being generated. Algorithms now have much more data to work with, and that alone means they can perform much better. In parallel, however, progress was made in domains such as image recognition: adjustments in neural networks, brought about by vibrant communities, have boosted the accuracy of the algorithms. ImageNet is a good example of this.

Much of what's being sold as "AI" today is snake oil. It does not and cannot work. Recently, Arvind Narayanan, a Princeton Professor, made waves by calling this out.
AI is continuing to make inroads in new domains at a breakneck pace. In 2019 alone, we've seen great progress in domains such as natural language processing, games, and common sense reasoning, to name but a few. New achievements in result quality and execution speed have been made almost monthly. The amount of resources dedicated is staggering, and research is progressing faster than ever. So, should we all be preparing for the brave new AI world? Well, maybe not so fast.
The problem with the AI frenzy is the divide between the haves and the have nots is widening. And not just because of the resources and expertise the big players have. It's a self-reinforcing loop of sorts: Being data-driven, designing and producing data-driven products means these products not only can have an edge, but they also bring in more data as they operate
As there is an evolutionary link connecting data and AI, more data is used to develop better AI, leading to better products, more data, and so on. An archetypal and widely recognized example of this is Facebook, but it's not the only one. When the likes of the Economist are calling for a new approach to antitrust rules for the data economy, this should be a cause for concern.
Data, however, is just one part of the AI equation. The other part is hardware. Without the tremendous progress in hardware, the 2010s have seen, AI would not be possible. Access to the compute power needed to process the massive amounts of data needed for machine learning used to be a privilege reserved for the select few. 
While the kind of hardware that Big Tech has access to remains beyond comprehension for most, democratization of sorts seems to have transcribed. The combination of cloud, with its on-demand access to processing power, and specialized hardware for AI workloads, has made AI chips accessible to more organizations than ever, assuming they can afford it.

Cerebras's "Wafer-Scale Engine" takes up almost all the area of a 12-inch silicon wafer, making it 57 times the size of Nvidia's largest graphics processing unit.
Cerebras Systems.
The big innovator, and winner, in the 2010s AI hardware was NVIDIA. The company that most people came to know as a maker of GPUs, specialized hardware typically used by gamers for fast graphics rendering, has reinvented itself as an AI superpower. The architecture of GPUs, it turns out, is very well suited to running AI workloads. 
Intel was becoming complacent in its dominance of traditional CPU hardware, and other GPU makers failed to execute, so NVIDIA rose to become the leader in AI hardware. That, however, is not set in stone, and the hardware space is already seeing rapid innovation.
While NVIDIA is dominating AI hardware and has built a software ecosystem around it too, waves of disruption are hitting the AI chip market. Just a few days before the closing of the 2010s, Intel stroke back by acquiring Habana Labs. Habana Labs is one of many startups in the AI chip market, looking to come up with new designs, built from the ground up to accommodate AI workloads. 
Even though for many Habana Labs is an unknown, its chips are already used in production by the likes of cloud vendors and autonomous vehicle makers. GraphCore, which became the first AI chip unicorn in late 2018, has recently announced its chips are now used in Microsoft Azure Cloud. Far from over, the AI chip race is only just beginning.

1. The Future is Graph, Knowledge Graph

Up until the beginning of the 2010s, the world was mostly running on relational databases and spreadsheets. To a large extent, it still does. But if the 2010s brought the first traces of dissent in the monoculture of tabular data structures, the 2020s will bring the final nail in the coffin. The NoSQL wave of databases has largely succeeded in getting developers, administrators, CIOs, CTOs, and business people out of their comfort zone, and instilled the "best tool for the job" mindset. 
Polyglot persistence, as is the lingo for using data models and data management interchangeably depending on the task at hand, is becoming the new normal. After relational, key-value, document, columnar, and time-series databases, the latest link in this evolutionary proliferation of data structures is graph. Graph databases and knowledge graphs have been making waves and being included in hype cycles for the last couple of years. 
While it's understandable why many people tend to think of graph as a new technology, the truth is this technology is at least 20 years old. And it has been largely initiated by none other than Tim Berners Lee, who is also credited as the inventor of the web, in 2001 with the publication of his Semantic Web manifesto in the Scientific American. Lee also coined the term Giant Global Graph, to describe the next stage in the evolution of the web.
Having been into this technology since the early 2000s, it's exhilarating to see it getting steam with technical progress, funding, and use cases piling up to a snowball effect. It is also amusing to see graph-washing beginning to commence. In essence, progress in graph is happening along the trajectory of progress in machine learning. 
It's not so much that there was a major breakthrough in the technology that made it feasible, but more about the right conditions that made it boom. Many of the concepts, formats, standards, and technology enabling graph databases and knowledge graphs to flourish today have been developed over more than 20 years. What has brought on the perfect graph storm is a combination of factors.

Google, NASA, and leading organizations from every domain are using knowledge graphs to manage and leverage vast amounts of data
Image: Google
Like AI, the data explosion has contributed to bringing graph in the fore. Now that Big is no longer a qualifier for Data, because we have mastered the art of storing lots of it, the question really is how to get value out of data. Leveraging connections in data is a prominent way of getting value out of data and graph is the best way of leveraging connections. 
This is why graph databases excel in use cases that require finding connections in data, such as anti-fraud or master data management. This is why graph analytics, with algorithms such as centrality or PageRank that are based in accounting for nodes and edges, can offer valuable insights in connected datasets. As the terminology seems to still be in flux for many newcomers in this field, a short history lesson, and grounding in semantics, may be called for. 
Graph analytics such as PageRank can be applied to data stored in any back end. Graph databases are back ends designed to accommodate graph data structures, offering specialized query languages, APIs, and oftentimes storage structures. Knowledge graphs, on the other hand, are a specific subclass of graphs, also called semantic graphs, that come with metadata, schema, and global identifier capabilities.
Google has played a key role in the rise of graphs, and knowledge graphs. As the web itself is a prime use case for graphs, PageRank was born. As crawling and categorizing content on the web is a very hard problem to solve without semantics and metadata, Google embraced them, and coined the term Knowledge Graph, in 2012. This, and the widespread adoption of schema.org that came with it, marked the beginning of the meteoric rise of graph technology and knowledge graphs. 
Knowledge graphs can address key challenges such as data governance but ultimately, they can serve as the digital substrate to unify the philosophy of knowledge acquisition and organization with the practice of data management in the digital age. The NASAs and the Morgan Stanleys of the world are managing ontologies, and utilizing knowledge graphs.
Graphs and knowledge graphs cross-cut into AI, too. Much of the AI hardware and software for the 2020s utilizes graph data structures. A combination of bottom-up, pattern matching techniques with top-down, knowledge-based approaches is the most promising way for AI to continue to make progress. 
As Nathan Benaich, author of the State of AI Report put it, "Domain knowledge can effectively help a deep learning system bootstrap its knowledge, by encoding primitives instead of forcing the model to learn these from scratch." Knowledge graphs are the best technology we have for encoding domain knowledge, and the world's most comprehensive knowledge base -- the web -- already functions as such.
Knowledge Graph is a technology that enables other technologies to accelerate their growth, and it also enables humans to take stock of their own knowledge. This is why the future is Knowledge Graph.

To infinity and beyond

Looking back, it becomes clear how far we have come in the relatively short span of the last decade. Counter-intuitive as this may seem, however, we are not certain this is a good thing. Somewhere along the way, technological progress left human ability to monitor, comprehend and digest technology in the dust. In the dawn of this new decade, we seem to be engrossed in the never-ending race for more: More data, more processing power, more technology.
The belief that more equals better seems to be firmly ingrained in most of us. And the signs of what's coming seem to tell the story of not just more, but immeasurably more. Quantum computing is progressing in leaps, promising to unlock compute power beyond our wildest imagination. DNA storage seems set to do the same for storage. More data and compute than we would know what to do with. To infinity and beyond. But what for, and for whom?
Is this technology making us happier, and bringing us closer, or is it alienating and distressing us? Where are all the huge productivity gains going? Who is in control, who gets to call the shots, and why?
AI, for example, is already being used to make critical decisions. Who gets to build those systems, on what data, and according to whose ethics and criteria? Should society as a whole have some sort of control over it? How could society even dream of controlling a technology it hardly understands, and in what way? Are we sure more technology is the solution to technologically induced issues? What is a moral compass for the 21st century?
For the time being, these are questions few people are prepared to tackle. But if the 2020s develop on the trajectory they are set on, more and more of us will have to face those questions head-on.

https://zd.net/37chjoe




Knowledge graph evolution: Platforms that speak your language

Knowledge graphs are among the most important technologies for the 2020s. Here is how they are evolving, with vendors and standard bodies listening, and platforms becoming fluent in many query languages

By for Big on Data



This may come as a shock if you've first encountered knowledge graphs in Gartner's hype cycles and trends, or in the extensive coverage they are getting lately. But here it is: Knowledge graph technology is about 20 years old. This, however, does not mean it's stagnating -- on the contrary. 



The 20-year old hype

First, let's quickly recap those 20 years of history. What we call Knowledge Graphs today has been largely initiated by none other than Tim Berners-Lee in 2001. Berners-Lee, who is also credited as the inventor of the web, published his Semantic Web manifesto in the Scientific American in 2001. The core concepts for Knowledge Graphs have been laid there.
The Semantic Web manifesto was in many ways ahead of its time. Looking back today, we can see some parts of it going strong, while others have faded. Building on a foundation of standards for interoperability, such as Unicode, URIs, and RDF, the core of the vision has always been semantics: instilling meaning in web content.
The Semantic Web technology stack, back in 2001. While its age shows, and some parts have become obsolete, others form the foundations of one of today's most hyped technologies: Knowledge Graphs


The Semantic Web got a bad name for being academic, while some technical choices such as XML did not quite work out. The thing is, however, that crawling and categorizing content on the web is a very hard problem to solve without semantics and metadata. This is why Google adopted the technology in 2010, by acquiring MetaWeb.
In 2012, the term Knowledge Graph was introduced. A very successful rebranding indeed, and that's not all we have Google to thank for. Google employs key people in the domain and is the driving force behind schema.org. Schema.org is the core of Google's knowledge graph. It is, unsurprisingly, a schema.
Knowledge graphs and schemas are foundationally bound. While not all knowledge graphs are as big as Google's, every one of them is based on a schema. Knowledge graph neophytes do not always realize this, but whether it's implicit or explicit, there's always a schema. Which brings us to the point.

Knowledge graphs and graph databases









Knowledge graphs can be stored in any back end, from files to relational databases or document stores. But since they are, well, graphs, it does make sense to store them in a graph database. This greatly facilitates storage and retrieval, as graph databases offer specialized structures, APIs, and query languages tailored for graphs.
In addition, many graph databases today offer a lot more than just a store for data. They come packaged with algorithms for graph analytics, visualization capabilities, machine learning features, and development environments. They have essentially grown from databases to platforms. But there is further nuance here. 
Graph databases come in two main flavors, depending on which graph model they support: Property graph and RDF. In general, RDF graph databases emphasize semantics and interoperability, while property graph databases emphasize ease of use and performance.





Graph databases come in 2 main flavors, depending on which graph model they support: labeled property graph (LPG) and RDF. Image: The Year of the Graph

When it comes to knowledge graphs, RDF graph databases are a natural match. It's not impossible to build knowledge graphs on top of property graph databases. Usually, however, this results in having to learn knowledge management fundamentals the hard way, and re-implement relevant features. While lessons don't come for free, building on platforms centered around knowledge management helps.
Property graphs and RDF graphs are not that different conceptually. Having interoperability between them would be both possible and desirable. This is why in March 2019 a W3C workshop on web standardization for graph data took place, as the first step towards standardization in the graph database world.
A key element to bridge the gap is something called RDF* (RDF star). RDF* is a proposal to standardize a modeling construct for RDF graphs, namely the addition of properties to edges. Although this is possible in RDF, there is no standard way of doing it. Standardizing it would not only help interoperability with property graphs but also interoperability among RDF graphs.

From secret handshakes to RDF stars

As Steve Sarsfield, VP of Product in Cambridge Semantics put it, before RDF*, if people wanted to use edge properties in RDF graphs, they had to rely on secret handshakes. This is not ideal, especially considering one of the key advantages of the RDF stack is standardization and interoperability.
In the wake of the W3C initiative, a couple of RDF graph database vendors went ahead and implemented RDF*. Cambridge Semantics is one of them. Its AnzoGraph database supports RDF*, as well as SPARQL*. SPARQL is the standard query language for RDF, and SPARQL* is its extension that works with RDF*.
Cambridge Semantics recently unveiled AnzoGraph DB Version 2, and when discussing the release with Sarsfield, we wondered what their experience from the field has been. Are people asking for RDF*, has it helped adoption? Bridging the gap with property graphs has enabled AnzoGraph to get an implementation of Cypher, the most popular language for querying property graphs, underway.
Sarsfield noted that it's still relatively early days for knowledge graph adoption. As such, many of the organizations that use AnzoGraph tend to have highly skilled people on board. For them, switching between data models and query languages is not much of an issue. For mainstream adoption, however, this is important.
Stardog is another RDF graph database vendor that has implemented RDF*. Mike Grove, Stardog co-founder and VP Engineering, said this has been in the works for a while, and they are very excited about it. Stardog started working on the plumbing as part of the Stardog 7 development effort, and they were very happy to be able to ship the feature.
Regarding its reception, Grove noted that what people wanted was a more user-friendly way to have edge properties: "Neo4j obviously got this right. RDF* does a fantastic job of bringing the same ease of use to semantic graphs." He went on to add that customers are excited, and many are already working on integrating it into their applications.
Technically, RDF* and SPARQL* are not yet standardized. Both have been introduced by Olaf Hartig, a researcher at Linköping University. When inquiring about their status, Hartig noted that while there have been delays, he hopes the standardization process will pick up speed soon.

For knowledge graph platforms, too, GraphQL is a plus

Both Sarsfield and Grove noted that they expect RDF* to boost knowledge graph adoption. Implementation is key, and having early adopters and real-world usage may also catalyze the standardization process. Sarsfield and Grove expressed their support for the process, as well as the need to get the word out.
RDF* can make a difference, but it's not the only thing going on in the knowledge graph world. As knowledge graphs entail several layers and can be a central piece of infrastructure for organizations, graph databases are growing into platforms.
AnzoGraph started as part of the Anzo platform before becoming a product in its own right. Stardog also touts its product as a platform, emphasizing features such as visualization and virtualization built around the graph database core.





GraphQL has benefits, and a new breed of approaches for using it as an access layer for databases is emerging. Image: Nordic APIs

Another RDF graph database vendor, Ontotext, recently announced a new version of its own platform. An interesting feature that Stardog's and Ontotext's platforms share is support for GraphQL. Unfortunately, GraphQL's name does not do it justice. As if there was not enough confusion already regarding graph: GraphQL is not a graph query language.
GraphQL is a replacement for REST APIs. Despite the misnomer, it's very useful, and its popularity among developers is growing. This is why more and more databases are adding support for GraphQL, with names such as MongoDB joining the GraphQL wave. Graph databases are no exception. Stardog has had it since 2017, Ontotext is in the process of adding it.
As Stardog put it, more developers know and are learning GraphQL than all the graph query languages combined. Ontotext on its part put together a rather elaborate post on the use of GraphQL in its platform. Whichever way you approach it, however, GraphQL makes lots of sense for accessing services built around database platforms.

GraphQL plus variants

Stardog reports GraphQL success within its customer base. Grove mentioned that one of the big Silicon Valley tech companies exclusively uses GraphQL to interact with Stardog. Both Grove and Jem Rayfield, Ontotext's Chief Architect, agree that GraphQL can work well in some cases, but by its very design, the expressiveness of GraphQL is quite limited.
Most people who don't know GraphQL assume it's a graph database query language. Most people who know GraphQL wonder how a graph database can be powered by it. This statement comes from Manish Jain, the CEO and founder of Dgraph. Dgraph is a graph database powered by GraphQL -- or something like it.
GraphQL+ is a derivative of GraphQL, developed and used exclusively by DGraph until today. In a 2019 interview with ZDNet, Jain expressed no interest in standardization for GraphQL+. No other vendor we know of has expressed interest in adopting GraphQL+ either. But that's not all there is to GraphQL for graph databases.
Most approaches are about what GraphQL can do for knowledge graphs. But to close the loop with the Semantic Web underpinning of knowledge graphs, here's an idea: What if GraphQL resources were annotated with URIs?
URIs are global identifiers, which can denote concepts from shared vocabularies, such as schema.org or other ontologies. This seems like a natural fit, and one that both Grove and Rayfield agree has potential. There is another working group set up to align RDF and GraphQL, although it does not look like it's moving very fast.

Knowledge graphs in the 2020s: We speak your language

It seems we are moving towards a new status quo. If NoSQL stands for Not Only SQL, we could call this NoSPARQL -- Not Only SPARQL. SPARQL remains the language of choice for taking full advantage of knowledge graph capabilities. It also doubles as an API, its expressiveness is beyond what GraphQL can attain, and SPARQL's federated query and data integration capabilities are unique.
But vendors seem set to meet users where they are, be it GraphQL or any other language. Even SQL. As Stardog's Grove put it: "We've always strived to bring our technology to the users. GraphQL was a step in that plan. Supporting SQL is the next step in that journey, not because SQL is better than GraphQL, but because of what that support enables."






Graph database vendors and standard bodies are listening to the market, and platforms fluent in many query languages are evolving
Getty Images/iStockphoto

SQL enables existing tooling to work on top of graph databases, making them accessible to a wider audience. Stardog is not the first graph database platform to have added an SQL connectivity layer. Cambridge Semantics also offers a connectivity layer for Tableau. More graph databases support SQL, and there is an ongoing standardization effort to add graph extensions to SQL itself.
Eventually, even natural language support could be an option. "No matter how you feel about SQL, SPARQL, GraphQL, or any other query syntax/language, natural language is just better. Why ask someone to learn an esoteric syntax when they can just simply type?" said Grove.
Grove mentioned Stardog will be launching a natural language interface to the knowledge graph. A pipedream? This may not be too far off. There is ongoing research for natural language interfaces for databases. And, to add to this, there are also existing integrations for accessing databases via voice assistants. So, you can see where this is going.
We don't know whether conversational knowledge graphs are something everyone would be comfortable with. What we do know is that more options is a good thing, and exciting times are ahead. Stay tuned as we keep exploring the years of the graph.

https://zd.net/39b3K9Y