Академический Документы
Профессиональный Документы
Культура Документы
Zeitgeist Australia
To BHMA Newspaper
Wired Magazine
Wired Magazine
How can I integrate all these data sources to prepare input for
analysis? What about data cleaning and transformations?
E.g. use Hadoop to collect in one system
Data Modeling,
Relational Algebra,
Query Languages,
Query Execution
Databases Data Modeling (1)
Model: a set of constructs and operations to describe a
real-world situation (mathematics, physics, chemistry
even art)
Data Model: a model to describe
data
data relationships
data semantics
data constraints
Conceptual/Logical Level:
How data is described using
some data model
SELECT d.account-number
FROM customer as c, depositor as d
WHERE c.customer-id = d.customer-id AND
c.customer-city=Palo Alto
account-number account-number
customer-city=Palo Alto
account-number
nested-loop join
merge join
linear scan
hash join
sorted scan customer-city=Palo Alto depositor
indexed scan
customer
Transaction Processing,
Concurrency & Recovery,
Parallel & Distributed Databases,
Main-Memory Databases
Databases Transactions, Definition (1)
A transaction is a unit of program execution a single
logical function that accesses and possibly updates
various data items.
Example: transfer 50 from account A to B:
1. read(A)
2. A := A 50
3. write(A)
4. read(B)
5. B := B + 50
6. write(B)
SSN Martha
219-19-9192 1
Mary
310-11-2212 2
492-00-3964 Salary
3 29.000
36.000
38.000
9/24/2016 An Introduction to Modern Data Management 88
Column-oriented Databases (1)
Data Warehousing,
OLAP, Ad Hoc Data Analysis,
Data Mining
Data Warehouses Definitions
A single version of the truth [Bill Inmon].
A single, complete and consistent store of data
obtained from different sources, made available to end
users in a what they can understand and use in a
business context [Barry Devlin].
A data warehouse is a subject-oriented, integrated,
time-variant, and nonvolatile collection of data in
support of managements decision-making process [Bill
Inmon].
Extract Query/Reporting
Transform Serve
Load
Operational DBs Refresh
Data Mining
Database
Statistics
Technology
Machine
Data Mining Visualization
Learning
Information Other
Science Disciplines
Wired Magazine
To BHMA Newspaper
Register Query
Results
Stream Query
Processor
Replication
Failover
Load Balancing
parallel processing
fault-tolerance
(http://www.wired.com/science/discoveries/magazine/16-07/pb_theory)