Google Website Translator Gadget

Wednesday, April 20, 2011

to translate



#s1#
#abstract1.txt#


The simulation of large groups (crowds) of individuals is a complex and
challenging task. It requires creating an adequate model that takes
into account psychological features of individuals, as well as
developing and implementing efficient computation and communication
strategies in order to manage the immense computational workload
implied by the numerous interactions within a crowd. This paper
develops a novel model for a realistic, real-time simulation of large
and dense crowds, focusing on evacuation scenarios and the modeling of
panic situations. Our approach ensures that both global navigation and
local motion are modeled close to reality, and the user can flexibly
change both the simulation environment and parameters at runtime.
Because of the high computation intensity of the model, we implement
the simulation in a distributed manner on multiple server machines,
using the RTF (Real-Time Framework) middleware. We implement state
replication as an alternative to the traditional state distribution via
zoning. We show that RTF enables a high-level development of
distributed simulations, and supports an efficient runtime execution on
multiple servers.

#e1#

#s1#
#abstract2.txt#


A Hardware Transactional Memory (HTM) aids the construction of
lock-free regions within applications with fewer concerns about
correctness and potentially greater performance through optimistic
concurrency. Object-aware hardware adds a level of indirection to
memory accesses, memory addresses become a combination of the object
being accessed and the offset within it. The hardware maintains a
mapping from objects to memory locations just as a mapping from virtual
to real memory is handled through page tables. In a scalable
object-aware system the directories are addressed by objects
identifiers. In this paper we extend a scalable object-aware memory
system to implement a HTM. Our object-aware protocol permits locks on
directories to be avoided for objects only read during a transaction.
Working at the granularity of an object allows entries within the
directories to be associated with multiple cache lines, as opposed to
one, and reduce the amount of network traffic. Finally, our commit
protocol dispenses with the need for a centrally controlled transaction
ID order.

#e1#

#s1#
#abstract10.txt#


A digital architecture for the emulation of dynamic bacterial quorum
sensing is presented. The architecture is completely scalable, with all
parameters stored in artificial bacteria. It shows that self-organizing
principles can be emulated in an electronic circuit, to build a
parallel system capable of processing information, similar to
cell-based structure of biological creatures. A mathematical model of
artificial bacteria and their behavior reflecting the self-organization
of the system are presented. The bacteria implemented on an 8-bit
microcontroller; and a framework with CPLDs to build the hardware
platform where the bacterial population increases are shown. Finally,
simulation results show the ability of the system to keep working after
physical damage, just as its biological counterpart.

#e1#

#s1#
#abstract11.txt#


ZnO nanorod arrays are grown on a-plane GaN template/r-plane sapphire
substrates by hydrothermal technique. Aqueous solutions of zinc nitrate
hexahydrate and hexamethylenetetramine were employed as growth
precursors. Electron microscopy and X-ray diffraction measurements were
carried out for morphology, phase and growth orientation analysis.
Single crystalline nanorods were found to have off-normal growth and
showed well-defined in-plane epitaxial relationship with the GaN
template. The #0001# axis of the ZnO nanorods were observed to be
parallel to the #101 macr 0# of the a-plane GaN layer. Optical property
of the as-grown ZnO nanorods was analyzed by room temperature
photoluminescence measurements. [All rights reserved Elsevier].

#e1#

#s1#
#abstract12.txt#


We report high aspect ratio nanochannel fabrication in glass using
single-shot femtosecond Bessel beams of sub-3 mu J pulse energies at
800 nm. We obtain near-parallel nanochannels with diameters in the
range 200-800 nm, and aspect ratios that can exceed 100. An array of
230 nm diameter channels with 1.6 mu m pitch illustrates the
reproducibility of this approach and the potential for writing periodic
structures. We also report proof-of-principle machining of a
through-channel of 400 nm diameter in a 43 mu m thick membrane. These
results represent a significant advance of femtosecond laser ablation
technology into the nanometric regime.

#e1#

#s1#
#abstract13.txt#


The following topics are dealt with: support vector machine; fuzzy
self-learning control; induction machines; online insurance business
process; marketing decision system; industrial control network;
integrated financial system; GA-BP neural network; sintering; image
fusion; speed control system; e-commerce tourism; travel market; UHF
RFID reader; GSM/CDMA repeater; single-patch antenna; pendulum
jawbreaker; agriculture expert system; container ship stowage system;
gyroscope; industrial robots; ZigBee wireless sensor networks; merger
and acquisition; garment industry; parallel batch scheduling problem;
logistics distribution; transmission line project; telemetry;
telecontrol;electric locomotives; talent cultivation; video encoding;
arc furnace; vehicle rollover dynamic; fault diagnosis; hydropower
dispatching; wireless ad hoc network; edge detection; campus network
information security; radar system; network-on-chips; ZVS-SEPIC
converter; decision support system; routing protocol; risk management;
network surveillance; body area communication channel; loudspeaker;
reconfigurable instrument; farm load balancing; rapid extended
manufacturing system; sensor signal feature extraction; product data
management;optical fiber communication; information sharing ;truck
loading balance monitoring system and wood constituents extraction.

#e1#

#s1#
#abstract14.txt#


In this paper, we analyze how humans can solve problems quickly by the
process of solving the problem of convex hull. Then we present the M2M
(Macro to Micro) data structure which maintains a finite set of n
points in the plane under insertion and deletion of points in amortized
O(1) time per operation and O(n) space usage. In addition, as the
insert operation of each point is independent, the algorithm has high
parallelism. And because the insert operations will not cause the
imbalance of the tree structure, the M2M data structure is dynamic.
Moreover, it can be shared by all the algorithms based on M2M model
which greatly improves the efficiency when a variety of algorithms work
simultaneity, such as in image processing and pattern recognition. In
all, the M2M model points out a general pattern for designing high
parallel algorithms, an efficient strategy for solving
multi-operational problems and a new approach for computer to stimulate
the thinking pattern of human beings.

#e1#

#s1#
#abstract15.txt#


During the last decade, the use of parallel and distributed systems has
become more common. In these systems, a huge chunk of data or
computation is distributed among many systems in order to obtain better
performance. Dividing data is one of the challenges in this type of
systems. Divisible Load Theory (DLT) is one of the old proposed method
for scheduling data distribution in parallel or distributed systems. In
many researches carried out in this field, it was assumed that all the
processors are dedicated for grid system but that is not always true in
real systems. The limited number of studies which attended to this
reality assumed that systems are homogeneous and presented some
algorithms or closed-formulas for scheduling jobs in a System with
Different Processors Availability Time (SDPAT). One of the examples of
arbitrary processor's available time is the last installment in
multi-installment systems. In this type of systems, usually two
different methods are used for scheduling internal and the last
installments. In this article, we will present a method for scheduling
the last installment in a heterogeneous multi-installment system.

#e1#

#s1#
#abstract16.txt#


With the development of computer technology and calculation methods,
the finite element method to analyze wheel/rail rolling contact became
popular. For the increasing requirement of calculation scale and
computing accuracy, the parallel computing method and parallel
computing environment become an effective way to solve this problem.
The parallel computing methods of contact problem is analyzed firstly.
Then, the contact algorithm and parallel computing of ABAQUS is
introduced. And the parallel computing environment using MPI in ABAQUS
is put forward. On the basis of cluster, some different finite element
model is solved by implicit and explicit solution. It is found that the
mesh size of wheel/rail contact field is refined to 0.75mm in order to
ensure accuracy for engineering. At last, the parallel computing for
the contact problem of wheel/rail is discussed using the speedup and
efficiency.

#e1#

#s1#
#abstract17.txt#


Image can be parsed into two main categories of representation:
structure shape and region texture. In this paper, the parallel
multi-regions image restoration system is proposed in order to ensure
the real-time application. This system is implemented on a dedicated
cluster. With different number of regions, PMR system is executed on
the same original image in order to obtain the best restoration result.
Experiments have demonstrated that our PMR system can improve the
restoration efficiency significantly and achieve optimal restoration.

#e1#

#s1#
#abstract18.txt#


Based on the research on the existing grid and parallel computing
technologies, this paper proposes a grid-based parallel computing
platform design scheme, and implements the key modules in it. The test
results show that the platform acceleration effect is very significant,
and this platform is simple, fast and easy to implement.

#e1#

#s1#
#abstract19.txt#


Because various machine translation (MT) systems adopt different
language rules and statistical methods, their performances are
different. To combine different MT systems and make full use of their
advantages so as to improve comprehensive performance of the
combination MT system, has become an important development trend in the
research field. This paper divides the current combination methods into
serial, parallel and hybrid combinations. A detailed introduction and
comparison of the three methods are given. In the last, authors briefly
summarize different methods and present the better method for
multi-systems combination.

#e1#

#s1#
#abstract20.txt#


Compared with the traditional distributed-memory High Performance
Computing (DMHPC), the shared-memory High Performance Computing (SMHPC)
not only saves data-transporting time, but also simplifies
computation-process. Furthermore, it predominates in solving special
problems. The performance of SMHPC has been studied by CFD application.
The results indicated that the CFD problems with small size of grids
could not reflect the advantage of SMHPC. As the advantage of the
performance of CPU, serial run was faster by using DMHPC than SMHPC,
however, with the increasing of processors, the advantage of SMHPC
network became the dominant factor. SMHPC parallel run came out to be
much faster. The largest speedup value of 11.78 achieved here on the
parallelization using SMHPC with 30 processors was encouraging. SMHPC
was a very easy-to-use and efficient parallel computing tool for the
CFD problems.

#e1#

#s1#
#abstract21.txt#


The Quadrature Mirror Filter (QMF) basically is a parallel combination
of of a High Pass Filter (HPF) and Low Pass Filter (LPF), which
performs the action of frequency subdivision by splitting the signal
spectrum into two spectra. Thus QMF finds wide applications in many
signal processing tasks such as transmultiplexing, equalizing wireless
communication channels; sub-band coding of speech and image signals,
sub-band acoustic echo cancellation etc. Hence this paper is devoted to
the effective utilization of XSG for hardware implementation of QMF.
This assists the Xilinx IP Core generator to deliver the optimized
results for the implemented device within short duration of time. The
results obtained in the simulation of QMF using XSG and the
communication system built using QMF at transmitter and receiver side
indicate that the nearly `Perfect Reconstruction' is possible without
any effect of aliasing. The obtained Synthesis Report for implemented
QMF indicates that the utilization of silicon die area and power
dissipation, for Spartan 3E device, are sufficiently low. Also the
comparison between the results of XSG and Simulink depicts that it is
efficient to built QMF using XSG. The study reveals that QMF-DPCM
provides much better approach towards the communication system. DPCM
using QMF is introduced to pave the way for smooth transition from
conventional 4 kHz band telephone systems to 7 kHz wide-band systems.

#e1#

#s1#
#abstract100.txt#


In this paper, a lens distortion correction method for low-cost digital
camera is proposed. The distortion coefficient and distortion center
are estimated by using geometric invariants of perspective projection.
The geometric invariants, including cross ratio of collinear points and
straight-parallel-perpendicular lines, will be invariant in
transforming from the world coordinate system to image coordinate
system if there exists no distortion. We derive new distortion measure
that is based on these geometric properties and can be optimized with
nonlinear search technique. The method is easy to apply and lead to
robust results with moderate effort. We verify accuracy and efficiency
from experiments.

#e1#

#s1#
#abstract101.txt#


We use parallel weighted finite-state transducers to implement a
part-of-speech tagger, which obtains state-of-the-art accuracy when
used to tag the Europarl corpora for Finnish, Swedish and English. Our
system consists of a weighted lexicon and a guesser combined with a
bigram model factored into two weighted transducers. We use both lemmas
and tag sequences in the bigram model, which guarantees reliable bigram
estimates.

#e1#

#s1#
#abstract102.txt#


In the present paper we describe TectoMT, a multi-purpose open-source
NLP framework. It allows for fast and efficient development of NLP
applications by exploiting a wide range of software modules already
integrated in TectoMT, such as tools for sentence segmentation,
tokenization, morphological analysis, POS tagging, shallow and deep
syntax parsing, named entity recognition, anaphora resolution,
tree-to-tree translation, natural language generation, word-level
alignment of parallel corpora, and other tasks. One of the most complex
applications of TectoMT is the English-Czech machine translation system
with transfer on deep syntactic (tectogrammatical) layer. Several
modules are available also for other languages (German, Russian,
Arabic). Where possible, modules are implemented in a
language-independent way, so they can be reused in many applications.

#e1#

#s1#
#abstract103.txt#


Summary form only given. Lexical semantic resources are a key component
of many NLP systems, whose performance continues to be limited by the
"lexical bottleneck". Two large hand-constructed resources, WordNet and
FrameNet, differ in their theoretical foundations and their approaches
to the representation of word meaning. A core question that both
resources address is, how can regularities in the lexicon be discovered
and encoded in a way that allows both human annotators and machines to
better discriminate and interpret word meanings? WordNet organizes the
bulk of the English lexicon into a network (an acyclic graph) of word
form-meaning pairs that are interconnected via directed arcs that
express paradigmatic semantic relations. This classification largely
disregards syntagmatic properties such as argument selection for verbs.
However, a comparison with a syntax-based approach like Levin (1993)
reveals some overlap as well as systematic divergences that can be
straightforwardly ascribed to the different classification principles.
FrameNet's units are cognitive schemas (Frames), each characterized by
a set of lexemes from different parts of speech with Frame-specific
meanings (lexial units) and roles (Frame Elements). FrameNet also
encodes cross-frame relations that parallel the relations among
WordNet's synsets. Given the somewhat complementary nature of the two
resources, an alignment would have at least the following potential
advantages: (1) both sense inventories are checked and corrected where
necessary, and (2) FrameNet's coverage (lexical units per Frame) can be
increased by taking advantage of WordNet's class-based organization. A
number of automatic alignments have been attempted, with variations on
a few intuitively plausible algorithms. Often, the result is limited,
as implicit assumptions concerning the systematicity of WordNet's
encoding or the semantic correspondences across the resources are not
fully warranted. Thus, not all members of a synonym set or a
subsumption tree are necessarily Frame mates. We carry out a manual
alignment of selected word forms against tokens in the American
National Corpus that can serve as a basis for semi-automatic alignment.
This work addresses a persistent, unresolved question, namely, to what
extent can humans select, and agree on, the context-appropriate meaning
of a word with respect to a lexical resource? We discuss representative
cases, their challenges and solutions for alignment as well as initial
steps for semi-automatic alignment.

#e1#

#s1#
#abstract104.txt#


In new DSP applications, reconfigurable architectures have emerged to
provide a flexible, high-performance, high-speed and low-power
implementation platform for wireless embedded devices. Since some DSP
algorithms rely heavily on multiplication, there are still demands for
more efficient multiplication structures. In this study, two
reconfigurable recursive multipliers are presented. The authors'
architectures combine some of the flexibility of software with the high
performance of hardware through implementing different levels of
recursive multiplication schemes on a two-dimensional logarithmic
number system (2DLNS) processing structure. The data are split into a
number of smaller sections, where each section is converted to a
2-digit 2DLNS (2 bases) representation. The dynamic range reduction and
logarithmic characteristics of computing with two orthogonal base
exponents in this number system allows multiplication to be implemented
with simple parallel small adders. The authors' architectures are able
to perform single and double precision multiplications, as well as
fault tolerant and dual throughput single precision operations. The
implementations demonstrate the efficiency of 2DLNS in multiplication
intensive DSP applications and show outstanding results in terms of
operation delay and dynamic power consumption.

#e1#

#s1#
#abstract105.txt#


This study presents a fast algorithm and its very large scale
integration (VLSI) design to implement the variable block size motion
estimation. The fast algorithm is proposed with a hardware-oriented
concept for regular VLSI design. Simulations show that the proposed
algorithm can reduce about 90% motion searching time, whereas P

#e1#

#s1#
#abstract106.txt#


Data processing on the cloud is increasingly used for offering cost
effective services. In this paper, we present a method for resource
allocation for data processing services over the cloud taking into
account not just the processing power and memory requirements, but the
network speed, reliability and data throughput. We also present
algorithms for partitioning data, for doing parallel block data
transfer to achieve better throughput and allocated cloud resources. We
also present methods for optimal pricing and determination of Service
Level Agreements for a given data processing job. The usefulness of our
approach is shown through experiments performed under different
resource allocation conditions.

#e1#

#s1#
#abstract107.txt#


A composite service can be constructed with the arbitrary combination
of sequential, parallel, loop, and conditional structures. In this
paper, we propose a general solution to calculate the QoS for composite
services with complex structures. We also show QoS-based service
selection can be conducted based on the proposed QoS calculation
method. An application example is given to show the effectiveness of
the method.

#e1#

#s1#
#abstract108.txt#


Video streaming with HDTV or UHDV quality will be provided and widely
demanded in the future. However, the transmission bit-rate of
high-quality video streaming is quite large, so generated traffic flows
will cause link congestion. Therefore, when providing streaming
services of rich content, it is important to flatten the link
utilization, i.e., reduce the maximum link utilization. To achieve this
goal, parallel video streaming in which ISPs use multiple servers to
deliver rich content is effective. However, the effect of parallel
video streaming depends on the network topology and link capacities. In
this paper, we investigate the impact of network topologies on the
effect of parallel video streaming using 23 actual commercial ISP
networks, when optimally designing server locations and optimally
selecting servers.

#e1#

#s1#
#abstract109.txt#


Internet traffic measurement and analysis have been usually performed
on a high performance server that collects and examines packet or flow
traces. However, when we monitor a large volume of traffic data for
detailed statistics, a long-period or a large-scale network, it is not
easy to handle Tera or Peta-byte traffic data with a single server.
Common ways to reduce a large volume of continuously monitored traffic
data are packet sampling or flow aggregation that results in coarse
traffic statistics. As distributed parallel processing schemes have
been recently developed due to the cloud computing platform and the
cluster filesystem, they could be usefully applied to analyzing big
traffic data. Thus, in this paper, we propose an Internet flow analysis
method based on the MapReduce software framework of the cloud computing
platform for a large-scale network. From the experiments with an
open-source MapReduce system, Hadoop, we have verified that the
MapReduce-based flow analysis method improves the flow statistics
computation time by 72%, when compared with the popular flow data
processing tool, flow-tools, on a single host. In addition, we showed
that MapReduce-based programs complete the flow analysis job against a
single node failure.

#e1#

#s1#
#abstract110.txt#


Inspired by the biological immunology principle and using its property
of autonomy, learning and memory for reference, this paper focuses on
the fault-tolerant design method of array processors which has
real-time detection and dynamic configuration capabilities, to improve
the chip's reliability by ensuring that when fault occurs in one or
more of processor elements the chip can also work normally. This paper
discusses the heterogeneous features of array processors, the structure
of resource node and communication node. Research the fault-tolerance
strategy of array processor, and the immune response mechanism of array
processor, to achieve the real-time immune process of perception,
training, response and feedback. This paper focuses on researching the
structure and the algorithm of router switch unit with 90 mm technology
which has the dynamic fault-tolerant function. It has the guiding
significance on R & D of array processors for industry circles.

#e1#

#s1#
#abstract111.txt#


This paper deals with the improvement of free energy computing in
segmentation algorithm according to the analysis of Image Segmentation
Algorithm Based on Boltzmann Theory (ISABT) and a new one is proposed.
It is proved by comparative experiments that performance indexes of the
new algorithm are considerably improved. The parallel processing method
of the algorithm which is applied to multi-processors is presented for
solving the real-time problem, and then its feasibility is verified by
experiments. The experiment shows that the improved ISABT with rapid
parallel image processing greatly enhances the speed of segmentation
with its segmentation effect assured. Compared with the original
algorithm, the new one may reach a speed of 20ms/F for 480*320 image on
the processing platform including eight pieces of TMS320DM642, and it
basically meets the real-time requirement for robot visual processing.

#e1#

#s1#
#abstract112.txt#


In parallel database system, optimizing distribution of relations could
improve processing efficiency of multi-join queries greatly. The cost
of data communication is expensive in parallel system based on PC
clusters. This paper proposes distribution of relations algorithm to
select an appropriate data placement strategy for each relation, which
includes selection of distribution attributes and nodes. The algorithm
could make best use of intra-operator parallelism, independent
inter-operator parallelism and pipelined parallelism of PC clusters
system. At the same time, it could reduce additional communication cost
of data redistribution. The result of experiment indicates the
algorithm has good performance and contributes to promoting execution
efficiency of parallel multi-join queries.

#e1#

#s1#
#abstract113.txt#


In order to solve the compute-intensive character of image processing,
based on advantages of GPU parallel operation, parallel acceleration
processing technique is proposed for image. First, efficient
architecture of GPU is introduced that improves computational
efficiency, comparing with CPU. Then, Sobel edge detector and
homomorphic filtering, two representative image processing algorithms,
are embedded into GPU to validate the technique. Finally, tested image
data of different resolutions are used on CPU and GPU hardware platform
to compare computational efficiency of GPU and CPU. Experimental
results indicate that if data transfer time, between host memory and
device memory, is taken into account, speed of the two algorithms
implemented on GPU can be improved approximately 25 times and 49 times
as fast as CPU, respectively, and GPU is practical for image processing.

#e1#

#s1#
#abstract114.txt#


Microscopic traffic simulation is an effective way to analyze the
feature and behavior of transportation systems. However performing
microscopic simulations for large-scale networks remains a huge
computing problem. Intensive computation characteristics displayed in
traffic simulation can benefit from parallel, distributed processing.
Based on simplified models, a MPI based distributed parallel traffic
simulation system is implemented. And by applying monolateral
communication strategy, a high performance is acquired.

#e1#

#s1#
#abstract115.txt#


Cellular automata (CA) have been accepted as a good evolutionary
computational model for the simulation of complex physical systems.
They have been used for various applications, such as parallel
processing computations and number theory. In this paper, we studied
the applications of cellular automata for the modular multiplications;
we proposed two new architectures of multipliers based on cellular
automata over finite field GF(2/sup m/). Since they have regularity,
modularity and concurrency, they are suitable for VLSI implementation.
The proposed architectures can be easily implemented into the hardware
design of crypto-coprocessors.

#e1#

#s1#
#abstract116.txt#


This paper presents a high speed real-time target detection system for
non homogenous environment based on FPGA Technology. The system
implements a Backward Automatic Censored Ordered Statistics Detector
(B-A

#e1#

#s1#
#abstract117.txt#


Power consumption is becoming a concern in programmable logic design as
the size and performance of modern FPGAs increase. Data-parallel
applications can work on different parallelism level so as to achieve
different performance. This paper presents an investigation into the
best parallelism degree-operating frequency tradeoff in order to find
the optimum number of instances for each parallelizable task with
adequate operating frequency minimizing energy and power for a given
throughput constraint. Significant optimization can be achieved not
only in power and energy but also in terms of adaptation to environment
conditions such as variable data rate and scalability in the multimedia
applications.

#e1#

#s1#
#abstract118.txt#


This paper presents an image-based framework for measuring target
objects on an oblique plane by using a single CCD camera and two
parallel laser projectors beside the camera. Because of the alignment
of the laser beams which form in parallel with the optical axis,
projected spots in the image can be processed to establish
relationships between distance and pixel counts between the projected
spots in the image. Based on simple geometrical derivations without
complex image processing, the proposed approach can measure the
photographing distance, the distance between two arbitrary points on
the oblique surface, and the incline angle. Experiment results have
demonstrated the effectiveness of the proposed approach in measuring
distant objects on an oblique plane.

#e1#

#s1#
#abstract119.txt#


Oil-water two-phase flow is commonly encountered in industrial
processes of petroleum and chemical engineering. The process
parameters, such as velocity and phase concentration of the mixture are
of great importance in both scientific and engineering field. Cross
correlation technique is one of the methods of calculating the flow
velocity by deriving the transit time of the fluid flowing through a
pair of parallel mounted sensors on the target pipelines. The physical
meaning of the cross correlation velocity is still an open discussion
in the field of multi-phase flow measurement. In this research, a
Dual-plane Electrical Resistance Tomography (ERT) is adopted to cross
correlate the sensing data from two electrode planes. A method of
dynamically seeking suitable signal segment for cross-correlating
measured signals from two-phase flow is proposed. The results show that
the cross correlation velocity is a structural velocity within
two-phase flow, and the relationship of cross correlation velocity and
mixture velocity is affected by the water flowrate within the two-phase
flow.

#e1#

#s1#
#abstract120.txt#


Defects on fabricated semiconductor wafers tend to cluster in
distinguishable patterns. The ability to accurately identify these
patterns allows manufacturers to trace their root causes to a specific
process step or equipment. This paper deals with an algorithm that
automatically extracts defect clusters. The algorithm performs cluster
segmentation and detection by employing two separate and parallel
processes. This increases robustness while maintaining high accuracy
and speed of data processing. In this paper a new method that allows
users to select a tradeoff threshold point between the acceptable false
alarm and false rejection rates to suit their applications is
introduced.

#e1#

#s1#
#abstract121.txt#


Microprocessor clock rates-which for three decades doubled about every
18 months-have essentially stopped increasing. Instead the number of
processor cores (identical processing units capable of all usual
microprocessor functions) in a microprocessor is increasing
exponentially with time. In order to increase performance as the number
of cores increase, measurement analysis software will have to take
advantage of this parallelism. The purpose of this paper is to study
one example of a measurement analysis having serial dependencies among
the input data and to show that there is a practical parallel algorithm
despite the data dependencies within the measured time series. The
measurement analysis studied is transition localization in digital
signals. A parallel scan-type algorithm is presented. Results of
applying the parallel algorithm on both synthetic data and actual
measured data are presented, and the speedup obtained on an eight core
processor analyzed.

#e1#

#s1#
#abstract122.txt#


The panning sorter is introduced, offering a new approach to the design
of highly compact digital parallel-input sorters for low power 2D
applications, such as image processing and data switching, among
others. The result is believed to be the smallest sorter circuit for
this type of implementation, for the given time complexity.

#e1#

#s1#
#abstract123.txt#


Within the field of industrial image processing the use of colour
cameras becomes ever more common. Increasingly the established black
and white cameras are replaced by economical single-chip colour cameras
with Bayer pattern. The use of the additional colour information is
particularly important for recognition or inspection. Become
interesting however also for the geometric metrology, if measuring
tasks can be solved more robust or more exactly. However only few
suitable algorithms are available, in order to detect edges with the
necessary precision. All attempts require however additional
computation expenditure. On the basis of a new filter for edge
detection in colour images with subpixel precision, the implementation
on a pre-processing hardware platform is presented. Hardware
implemented filters offer the advantage that they can be used easily
with existing measuring software, since after the filtering a single
channel image is present, which unites the information of all colour
channels. Advanced field programmable gate arrays represent an ideal
platform for the parallel processing of multiple channels. The
effective implementation presupposes however a high programming
expenditure. On the example of the colour filter implementation,
arising problems are analyzed and the chosen solution method is
presented.

#e1#

#s1#
#abstract124.txt#


In this paper, we attempt to scale up the kd-tree indexing methods for
large-scale vision applications, e.g., indexing a large number of SIFT
features and other types of visual descriptors. To this end, we propose
an effective approach to generate near-optimal binary space
partitioning and need low time cost to access the nodes in the query
stage. First, we relax the coordinate-axis-alignment constraint in
partition axis selection used in conventional kd-trees, and form a
partition axis with the great variance by combining a few coordinate
axes in a binary manner for each node, which yields better space
partitioning and requires almost the same time cost to visit internal
nodes during the query stage thanks to cheap projection operations.
Then, we introduce a simple but very effective scheme to guarantee the
partition axis of each internal node is orthogonal to or parallel with
those of its ancestors, which leads to efficient distance computation
between a query point and the cell associated with each node and yields
fast priority search. Compared with the conventional kd-trees, our
approach takes a little more tree construction time, but obtains much
better nearest neighbor search performance. Experimental results on
large scale local patch indexing and image search with tiny images show
that our approach outperforms the state-of-the-art kd-tree based
indexing methods.

#e1#

#s1#
#abstract125.txt#


Graph-cuts optimization is prevalent in vision and graphics problems.
It is thus of great practical importance to parallelize the graph-cuts
optimization using today's ubiquitous multi-core machines. However, the
current best serial algorithm by Boykov and Kolmogorov (called the BK
algorithm) still has the superior empirical performance. It is
non-trivial to parallelize as expensive synchronization overhead easily
offsets the advantage of parallelism. In this paper, we propose a novel
adaptive bottom-up approach to parallelize the BK algorithm. We first
uniformly partition the graph into a number of regularly-shaped
disjoint subgraphs and process them in parallel, then we incrementally
merge the subgraphs in an adaptive way to obtain the global optimum.
The new algorithm has three benefits: 1) it is more cache-friendly
within smaller subgraphs; 2) it keeps balanced workloads among
computing cores; 3) it causes little overhead and is adaptable to the
number of available cores. Extensive experiments in common applications
such as 2D/3D image segmentations and 3D surface fitting demonstrate
the effectiveness of our approach.

#e1#

#s1#
#abstract126.txt#


We present an efficient and scalable technique for spatiotemporal
segmentation of long video sequences using a hierarchical graph-based
algorithm. We begin by over-segmenting a volumetric video graph into
space-time regions grouped by appearance. We then construct a "region
graph" over the obtained segmentation and iteratively repeat this
process over multiple levels to create a tree of spatio-temporal
segmentations. This hierarchical approach generates high quality
segmentations, which are temporally coherent with stable region
boundaries, and allows subsequent applications to choose from varying
levels of granularity. We further improve segmentation quality by using
dense optical flow to guide temporal connections in the initial graph.
We also propose two novel approaches to improve the scalability of our
technique: (a) a parallel out-of-core algorithm that can process
volumes much larger than an in-core algorithm, and (b) a clip-based
processing algorithm that divides the video into overlapping clips in
time, and segments them successively while enforcing consistency. We
demonstrate hierarchical segmentations on video shots as long as 40
seconds, and even support a streaming mode for arbitrarily long videos,
albeit without the ability to process them hierarchically.

#e1#

#s1#
#abstract127.txt#


Graph cuts methods are at the core of many state-of-the-art algorithms
in computer vision due to their efficiency in computing globally
optimal solutions. In this paper, we solve the maximum flow/minimum cut
problem in parallel by splitting the graph into multiple parts and
hence, further increase the computational efficacy of graph cuts.
Optimality of the solution is guaranteed by dual decomposition, or more
specifically, the solutions to the subproblems are constrained to be
equal on the overlap with dual variables. We demonstrate that our
approach both allows (i) faster processing on multi-core computers and
(ii) the capability to handle larger problems by splitting the graph
across multiple computers on a distributed network. Even though our
approach does not give a theoretical guarantee of speedup, an extensive
empirical evaluation on several applications with many different data
sets consistently shows good performance. An open source implementation
of the dual decomposition method is also made publicly available.

#e1#

#s1#
#abstract128.txt#


Tracking-by-detection is increasingly popular in order to tackle the
visual tracking problem. Existing adaptive methods suffer from the
drifting problem, since they rely on self-updates of an on-line
learning method. In contrast to previous work that tackled this problem
by employing semi-supervised or multiple-instance learning, we show
that augmenting an on-line learning method with complementary tracking
approaches can lead to more stable results. In particular, we use a
simple template model as a non-adaptive and thus stable component, a
novel optical-flow-based mean-shift tracker as highly adaptive element
and an on-line random forest as moderately adaptive appearance-based
learner. We combine these three trackers in a cascade. All of our
components run on GPUs or similar multi-core systems, which allows for
real-time performance. We show the superiority of our system over
current state-of-the-art tracking methods in several experiments on
publicly available data.

#e1#

#s1#
#abstract129.txt#


A catadioptric system consisting of a pinhole camera and two planar
mirrors is deeply investigated in this paper. The two mirrors combine
to form a corner and face-to-face with the pinhole. Their relative pose
is unknown. An object will be reflected in the mirror corner one-time
or multiple-times. Using the pinhole, we may take an image containing
the object and its reflections, i.e., simultaneously imaging multiple
views of an object by a single camera. We discovered that each 3D point
and its reflections lie on a circle. We call the point set composed of
a 3D point and its reflections, a Reflection Point Group (RPG), and the
circle related to a RPG is called a RPG circle. All RPG circles are
parallel to one another. Furthermore, each RPG can be partitioned into
two separate subgroups. Shape formed by all points in a subgroup is
invariant with respect to location of the 3D point. From these
geometric properties, two calibration approaches can be utilized: One
is based on parallel circles; the other is using 2D homographies among
invariant shapes. Experiments validate our approaches.

#e1#

#s1#
#abstract130.txt#


We present a flexible method for fusing information from optical and
range sensors based on an accelerated high-dimensional filtering
approach. Our system takes as input a sequence of monocular camera
images as well as a stream of sparse range measurements as obtained
from a laser or other sensor system. In contrast with existing
approaches, we do not assume that the depth and color data streams have
the same data rates or that the observed scene is fully static. Our
method produces a dense, high-resolution depth map of the scene,
automatically generating confidence values for every interpolated depth
point. We describe how to integrate priors on object motion and
appearance and how to achieve an efficient implementation using
parallel processing hardware such as GPUs.

#e1#

#s1#
#abstract131.txt#


We propose a novel tracking algorithm that can work robustly in a
challenging scenario such that several kinds of appearance and motion
changes of an object occur at the same time. Our algorithm is based on
a visual tracking decomposition scheme for the efficient design of
observation and motion models as well as trackers. In our scheme, the
observation model is decomposed into multiple basic observation models
that are constructed by sparse principal component analysis (SPCA) of a
set of feature templates. Each basic observation model covers a
specific appearance of the object. The motion model is also represented
by the combination of multiple basic motion models, each of which
covers a different type of motion. Then the multiple basic trackers are
designed by associating the basic observation models and the basic
motion models, so that each specific tracker takes charge of a certain
change in the object. All basic trackers are then integrated into one
compound tracker through an interactive Markov Chain Monte Carlo
(IMCMC) framework in which the basic trackers communicate with one
another interactively while run in parallel. By exchanging information
with others, each tracker further improves its performance, which
results in increasing the whole performance of tracking. Experimental
results show that our method tracks the object accurately and reliably
in realistic videos where the appearance and motion are drastically
changing over time.

#e1#

#s1#
#abstract132.txt#


This paper introduces an approach for enabling existing multi-view
stereo methods to operate on extremely large unstructured photo
collections. The main idea is to decompose the collection into a set of
overlapping sets of photos that can be processed in parallel, and to
merge the resulting reconstructions. This overlapping clustering
problem is formulated as a constrained optimization and solved
iteratively. The merging algorithm, designed to be parallel and
out-of-core, incorporates robust filtering steps to eliminate
low-quality reconstructions and enforce global visibility constraints.
The approach has been tested on several large datasets downloaded from
Flickr.com, including one with over ten thousand images, yielding a 3D
reconstruction with nearly thirty million points.

#e1#

#s1#
#abstract133.txt#


In this paper, we consider the problem of stereo matching using loopy
belief propagation. Unlike previous methods which focus on the original
spatial resolution, we hierarchically reduce the disparity search
range. By fixing the number of disparity levels on the original
resolution, our method solves the message updating problem in a time
linear in the number of pixels contained in the image and requires only
constant memory space. Specifically, for a 800 * 600 image with 300
disparities, our message updating method is about 30* faster (1.5
second) than standard method, and requires only about 0.6% memory (9
MB). Also, our algorithm lends itself to a parallel implementation. Our
GPU implementation (NVIDIA Geforce 8800GTX) is about 10* faster than
our CPU implementation. Given the trend toward higher-resolution
images, stereo matching using belief propagation with large number of
disparity levels as efficient as the small ones makes our method
future-proof. In addition to the computational and memory advantages,
our method is straightforward to implement.

#e1#

#s1#
#abstract134.txt#


We propose a novel method for automatic camera calibration and
foot-head homology estimation by observing persons standing at several
positions in the camera field of view. We demonstrate that human body
can be considered as a calibration target thus avoiding special
calibration objects or manually established fiducial points. First, by
assuming roughly parallel human poses we derive a new constraint which
allows to formulate the calibration of internal and external camera
parameters as a Quadratic Eigenvalue Problem. Secondly, we couple the
calibration with an improved effective integral contour based human
detector and use 3D projected models to capture a large variety of
person and camera mutual positions. The resulting camera
auto-calibration method is very robust and efficient, and thus well
suited for surveillance applications where the camera calibration
process cannot use special calibration targets and must be simple.

#e1#

#s1#
#abstract135.txt#


We implemented a hardware accelerated system for deep packet
inspection. The proposed system makes use of operation level and
connection level parallelism to achieve a maximum processing rate of
595.2Mb/s. The system supports 25 mandatory patterns running at 125 MHz
on a Xilinx Virtex 5 FPGA.

#e1#

#s1#
#abstract136.txt#


The MEMO

#e1#

#s1#
#abstract137.txt#


The book has 10 chapters which deals with the following topics:
foundations of algorithm engineering; modeling assessment; modeling
languages; mixed integer programming; constraint programming; parallel
computing; design scalable algorithms; grid computing; formal methods;
software engineering aspects; worst case analysis; realistic input
models; asymptotic performance; real architectures; memory hierarchy;
parallel algorithms for I/O efficiency; verification techniques; memory
management ; data structures; data libraries; Voronoi diagrams.

#e1#

#s1#
#abstract138.txt#


Many real-world applications involve storing and processing large
amounts of data. These data sets need to be either stored over the
memory hierarchy of one computer or distributed and processed over many
parallel computing devices or both. In fact, in many such applications,
choosing a realistic computation model proves to be a critical factor
in obtaining practically acceptable solutions. In this chapter, we
focus on realistic computation models that capture the running time of
algorithms involving large data sets on modern computers better than
the traditional RAM (and its parallel counterpart PRAM) model.

#e1#

#s1#
#abstract139.txt#


Server virtualization offers the ability to slice large, underutilized
physical servers into smaller, parallel virtual machines (VMs),
enabling diverse applications to run in isolated environments on a
shared hardware platform. Effective management of virtualized cloud
environments introduces new and unique challenges, such as efficient
CPU scheduling for virtual machines, effective allocation of virtual
machines to handle both CPU intensive and I/O intensive workloads.
Although a fair number of research projects have dedicated to
measuring, scheduling, and resource management of virtual machines,
there still lacks of in-depth understanding of the performance factors
that can impact the efficiency and effectiveness of resource
multiplexing and resource scheduling among virtual machines. In this
paper, we present our experimental study on the performance
interference in parallel processing of CPU and network intensive
workloads in the Xen Virtual Machine Monitors (VMMs). We conduct
extensive experiments to measure the performance interference among VMs
running network I/O workloads that are either CPU bound or network
bound. Based on our experiments and observations, we conclude with four
key findings that are critical to effective management of virtualized
cloud environments for both cloud service providers and cloud
consumers. First, running network-intensive workloads in isolated
environments on a shared hardware platform can lead to high overheads
due to extensive context switches and events in driver domain and VMM.
Second, co-locating CPU-intensive workloads in isolated environments on
a shared hardware platform can incur high CPU contention due to the
demand for fast memory pages exchanges in I/O channel. Third, running
CPU-intensive workloads and network-intensive workloads in conjunction
incurs the least resource contention, delivering higher aggregate
performance. Last but not the least, identifying factors that impact
the total demand of the exchanged memory pages is critical to the
in-depth understanding of the interference overheads in I/O channel in
the driver domain and VMM.

#e1#

#s1#
#abstract140.txt#


In computing clouds, it is desirable to avoid wasting resources as a
result of under-utilization and to avoid lengthy response times as a
result of over-utilization. In this paper, we propose a new approach
for dynamic autonomous resource management in computing clouds. The
main contribution of this work is two-fold. First, we adopt a
distributed architecture where resource management is decomposed into
independent tasks, each of which is performed by Autonomous Node Agents
that are tightly coupled with the physical machines in a data center.
Second, the Autonomous Node Agents carry out configurations in parallel
through Multiple Criteria Decision Analysis using the PROMETHEE method.
Simulation results show that the proposed approach is promising in
terms of scalability, feasibility and flexibility.

#e1#

#s1#
#abstract141.txt#


Most of the large-scale scientific experiments modeled as scientific
workflows produce a large amount of data and require workflow
parallelism to reduce workflow execution time. Some of the existing
Scientific Workflow Management Systems (SWfMS) explore parallelism
techniques - such as parameter sweep and data fragmentation. In those
systems, several computing resources are used to accomplish many
computational tasks in homogeneous environments, such as multiprocessor
machines or cluster systems. Cloud computing has become a popular high
performance computing model in which (virtualized) resources are
provided as services over the Web. Some scientists are starting to
adopt the cloud model in scientific domains and are moving their
scientific workflows (programs and data) from local environments to the
cloud. Nevertheless, it is still difficult for the scientist to express
a parallel computing paradigm for the workflow on the cloud. Capturing
distributed provenance data at the cloud is also an issue. Existing
approaches for executing scientific workflows using parallel processing
are mainly focused on homogeneous environments whereas, in the cloud,
the scientist has to manage new aspects such as initialization of
virtualized instances, scheduling over different cloud environments,
impact of data transferring and management of instance images. In this
paper we propose SciCumulus, a cloud middleware that explores parameter
sweep and data fragmentation parallelism in scientific workflow
activities (with provenance support). It works between the SWfMS and
the cloud. SciCumulus is designed considering cloud specificities. We
have evaluated our approach by executing simulated experiments to
analyze the overhead imposed by clouds on the workflow execution time.

#e1#

#s1#
#abstract142.txt#


The 1.5 m primary ZERODUR/sup reg / mirror of the solar telescope
GREGOR incorporates 420 pockets at the backside for active cooling to
avoid the thermal load impact of the sun deteriorating the observation.
This design is also under consideration for the 2 m Indian Solar
Telescope and for the 4.2 m European Solar Telescope (EST). The tip and
tilt M5 mirror of the European Extremely Large Telescope (E-ELT)
requires an even more demanding approach in light weighting. The
approximately 3 m * 2.4 m elliptical flat mirror is specified to a
weight of less than 500 kg. During the successful manufacturing of the
GREGOR light weighted mirror, SCHOTT developed a systematic approach
for processing such complex and long lead items which are capable for
being up-scaled to a dimension of 4 m. In parallel SCHOTT has tested
the machining of challenging aspect ratios of rib thickness and pocket
height to prove the machinability of the E-ELT M5 design suggestions.
The improved data on the bending strengths of ZERODUR/sup reg / enable
aggressive designs for light weighted 4 m class mirrors.

#e1#

#s1#
#abstract144.txt#


In order to demonstrate the Wedge as a parallel optic we have built a
flat panel multi-projector autostereoscopic 3D display. The bulk,
otherwise inherent in such displays, has been eliminated by use of a
wedge-shaped light-guide. Two 280-mm (11") diagonal VGA resolution
views were formed on a 406-mm (16") diagonal screen by pointing a pair
of LED-illuminated pico-projectors measuring 115 X 50 X 22 mm into the
20-mm thick input face of a 710-mm long acrylic wedge. Distortions
introduced by the Wedge were reduced to less than 2% by predistortion
algorithms and distortion between views was less than 0.4%. Between the
projectors was placed an infrared camera which imaged objects placed
directly in front of the 3D screen.

#e1#

#s1#
#abstract145.txt#


Real-time simulation of haptic interaction with deformable objects is
computationally demanding. In particular in finite-element (FE) based
analysis of such interactions, a large system of equations must be
solved at an update rate of 100-1,000 Hz for simulation fidelity and
stability. A new hardware-based parallel implementation of a
Preconditioned Conjugate Gradient (PCG) algorithm is proposed for
solving the linear systems of equations arising from FE-based
deformation models. Concurrent utilization of a large number of
fixed-point computing units on a Field-Programmable Gate Array (FPGA)
device yields a very fast solution to these equations. Quantization and
overflow errors in the fixed-point implementation of the iterative
solver are minimized through dynamic scaling and preconditioning.
Numerical accuracy of the solution, the architecture design, and issues
pertaining to the degree of parallelism and scalability of the
architecture are discussed in detail. The implementation of the solver
on an Altera EP3SE110 FPGA device has enabled real-time simulation of
three-dimensional linear elastic deformation models with 1,500 nodes at
an update rate of up to 2,500 Hz.

#e1#

#s1#
#abstract146.txt#


Large-scale erasure-coded storage systems have a serious performance
problem due to I/O congestion and disk media access congestion caused
by read-modify-write operations involved in small-write operations. All
the existing technologies based on the conventional disk can provide
very limited performance improvement. This paper presents a new Disk
Architecture with Composite Operation (DA

#e1#

#s1#
#abstract147.txt#


There are a large number of high-temperature sensing and frequency
control applications that can be addressed using acoustic wave devices
capable of operation at high-temperatures. For those applications, it
is important to characterize the acoustic properties of the
piezoelectric crystal used as substrate at elevated temperatures.
Langatate (LGT) is one of the crystals which allow the fabrication of
SAW devices at elevated temperatures. In a previous work, the authors
measured and discussed the LGT elastic constants up to 900 degrees C.
This paper reports the langatate complex dielectric permittivity and
conductivity from 25 to 900 degrees C. The constants were extracted
from impedance measurements of parallel-plate capacitors fabricated
with Pt/Rh/ZrO/sub 2/ electrodes on LGT wafers aligned along the X and
Z crystalline axes. The real permittivities, epsiv '/sub 11/ and epsiv
"/sub 33/, were found to change significantly in the range from 25 to
900 degrees C with a 38% increase and a 49% decrease of their room
temperature values, respectively. Thus, it is important to include the
extracted high temperature permittivities when designing LGT acoustic
wave devices and not simply to use extrapolated low temperature data.
Both LGT conductivity and imaginary permittivity are necessary to
quantify the electrical losses of sensors, signal-processing, and
frequency-control devices operating with this substrate at
high-temperatures.

#e1#

#s1#
#abstract148.txt#


Nowadays, software tools to test applications in several areas have
been developed by different research groups. However, to our knowledge,
no tools to make parameterized unity tests with images can be found in
the literature. In this work a prototype tool called ITVT (Image
Testing and Visualization Tool) has been developed, that allows for
performing parameterized unity tests and visualization of images. This
tool can be used in parallel with general-purpose programming,
environments. As an application of ITVT, several, implementations for
fingerprint segmentation were developed and tested. Differences were
found in the performance for the different algorithms.

#e1#

#s1#
#abstract149.txt#


In this paper we present a methodology and tool for rapid prototyping
of real time image processing applications. We describe our design flow
of multiprocessor system on chip (MPSoC) architectures based on
hardware/software components. This methodology provides automated
methods to specify, generate the hardware, software, and the
architectural interfaces between them. Our methodology starts from
system level specification of the application with parallel processes
described in C-code. The processes communicate through an abstract
channel called streams. We describe also the solution that we proposed
to synthesize a custom bus architecture for the reconfigurable
computing applications, which therefore allows minimizing hardware cost
in the FPGA. On the other hand, this methodology allows generating the
communication architecture based on the needs of the application.
Finally, we demonstrate the effectiveness of our approach through an
image processing application.

#e1#

#s1#
#abstract150.txt#


Radar observations from Marsis have demonstrated that Martian Polar
Layered Deposits (PLD's) are very transparent to radar waves. Thus, the
sounder is able to detect the presence of subsurface reflections in the
polar regions below the ice-rich layered deposits. The analysis of
radar data makes it possible to gain information about some physical
features of Mars surface. In this work an electromagnetic inversion
model is used to characterize the shallower structures. This approach
assumes that structure consists of layers with parallel plane
interfaces and that the electromagnetic properties of the first layer
are known a priori. Under these assumptions it is possible to estimate
the dielectric permittivity of the subsurface structure. The inversion
method has been tested in an area of South Pole and reconstruction
results are shown.

#e1#

#s1#
#abstract151.txt#


This paper proposes a self-diagnosable multi-agent system. A
self-diagnosable algorithm has been proposed for multi-agent systems.
This conventional algorithm makes every agent diagnose all other
agents, but it has a problem that the more agents are included in a
system, the more communications traffic increases. To solve the
problem, this paper proposes a framework of multi-agent system that
divides and deals with in parallel domains for mutual diagnosis to
mitigate the communications between agents. Our approach introduces
middle agents that do not interaction with basic client agents and we
have developed a dynamic highly structured system that dynamically
reconstitutes system configuration. The proposed method has been
applied to a multi-agent system that forms a circle autonomously.
Numerical experiments show that the proposed method needs much less
communications than the conventional one.

#e1#

#s1#
#abstract152.txt#


We propose an effective scheme based on redundant residue number
systems (RRNS) for radiation hardening of datapath in this paper. The
proposed hardening scheme employs redundant residues to mitigate soft
errors. We use the properties of RRNS, such as independence, parallel
and error correction, to establish the radiation hardening architecture
for datapath in radiation environments. In the architecture, all of the
residues can be processed independently, and most of soft errors in
datapath can be corrected with the redundant relationship of the
residues at correction module, which is allocated at the end of the
architecture. There is not any additional processing in the
architecture but correction operation and binary to residue
translation. Thus, the proposed scheme has high efficiency when the
processing steps of datapath are large. In order to provide protection
for the key module in the proposed scheme, correction module, some
traditional protection methods, including Guard-gate register triple
modular redundancy (G-RTMR), pipeline-level TMR, and module-level TMR,
are used to construct the hybrid protection schemes. In the case
studies, we find that the RRNS based scheme and the hybrid schemes can
reduce soft error rate from 10/sup -12/ to 10/sup -17/, 10/sup -18/,
10/sup -25/, and 10/sup -25/, respectively, only with 45.6-46.1% area
overheads and negligible latency overheads, when the pipeline stages of
datapath are 10/sup 4/ and the moduli set with correction capability of
one residue is adopted. The case studies also show that the proposed
schemes can achieve less area and latency overheads than that of
datapath without radiation hardening, because RRNS can reduce the
complexity of operations in datapath.

#e1#

#s1#
#abstract153.txt#


This paper proposes a new AC/DC power supply system with DGs using the
parallel processing method. The purpose of this system is to supply the
power without connecting to the utility gird. The DGs consist of
photovoltaic (PV) and wind generators (WG), and the main energy source
of this system is the DGs and the UPS battery. When the system does not
get enough energy from the DGs, and the power supply of the battery
runs out, this system connects to the utility grid and begins to charge
the battery and supply to the load. In this study, we focused on the AC
power supply and DC power supply from DGs. Then, the converter of DGs
uses two kinds of DC/DC type and DC/AC type. Also, the load uses two
kinds of DC load and AC load. The proposed the power system has been
operated with the actual loads of 20 kW in the campus of the Aichi
Institute of Technology.

#e1#

#s1#
#abstract154.txt#


Molecular docking technology is an important tool in molecular
recognition and structure prediction for the protein-protein complex.
In this work, based on the analysis of the widely used FFT-based
method, we proposed a GPU parallel molecular docking program. The new
docking program was applied to dock 5 protein-protein complexes of
enzyme/inhibitor type. Docking results indicate that the parallel codes
can make a good prediction of the given complex structures. Moreover,
the GPU docking method can achieve a speedup of more than 3 times.
Finally, we analyzed the influence of parameters for the docking
results and parallel speedup. The obtained high quality parallel
speedup and efficiency show the promising improvement of our method for
future applications on structure prediction of protein-protein complex.

#e1#

#s1#
#abstract155.txt#


Derived from the individual surface coil images, a new method based on
anisotropic diffusion for estimating the coil sensitivity is proposed
in this paper. When the coil sensitivity maps estimated by this method
are applied to reconstruct the magnetic resonance image from the
under-sampled images by parallel imaging reconstruction Sensitivity
Encoding method, the quality of the reconstructed image is improved. As
a result, without using the body coil for additional reference scans,
the proposed method is valuable for determining the coil sensitivity
maps alone from the surface coil images.

#e1#

#s1#
#abstract156.txt#


The Compute Unified Device Architecture (CUDA) is a new programming
platform making use of the unified shader design of the most current
Graphics Processing Units (GPUs) from NVIDIA. In this paper, we apply
this revolutionary new technology to implement the automatic time gain
compensation (ATGC) for medical ultrasound imaging. The parallel box
filtering method and general matrix computation algorithms are also
presented. This ATGC method achieves a frame rate of 125 fps for the
512*261 image, about 79 times faster than the CPU implementation.
Testing results from GPU and CPU are compared in terms of visual image
quality and program runtime with different image sizes.

#e1#

#s1#
#abstract157.txt#


A GPU framework for ultrasound color flow imaging (CFI) based on
auto-correlation is presented. The parallel CFI processing framework
implementation is mainly based on CUDA performance features, such as
the memory selection strategy, applicable thread structure and
high-throughput bandwidth. Parallel convolution algorithm and
multi-channel championship algorithm are proposed. This CFI method
achieves a frame rate of 300 fps from the Doppler signal, in which the
number of scan lines is 44, the number of samples along the axial line
is 510 and ensemble size is 16.

#e1#

#s1#
#abstract158.txt#


Ultrasonic tissue motion can be visualized in three steps: the gray
scale motion detection, the Unsteady Flow Line Integral Convolution
(UFLIC) algorithm to trace the velocity field, and display techniques
for both the global motion and the local radial/tangential velocity
components. Because of the large amount of data and high computational
requirement in UFLIC, it was difficult to meet real-time requirement on
CPU. In this paper a parallel algorithm based on the graphics
processing unit (GPU) was proposed to implement ultrasonic tissue
motion visualization, and it represented both the direction and
amplitude of the motion by texture and color. Furthermore, a method was
proposed to calculate the local radial/tangential velocity components.
Finally, we got a frame rate of about more than 300 fps with the vector
field size of 260*260.

#e1#

#s1#
#abstract159.txt#


We study the spectral features of the polarized fluorescence spectra of
normal and cancerous human breast tissues through continuous wavelet
transform, which clearly identifies distinguishing features between the
tissue types. After pinpointing these robust features in the wavelet
scalogram, we systematically study the autocorrelation property of the
wavelet coefficients of the fluorescence spectra, which is found to
differentiate normal and malignant tissues with high sensitivity. The
intensity difference of parallel and perpendicularly polarized
fluorescence spectra is subjected to investigation, since the same is
relatively free of the diffusive background.

#e1#

#s1#
#abstract160.txt#


MRI of the human heart without explicit cardiac synchronization
promises to extend the applicability of cardiac MR to a larger patient
population and potentially expand its diagnostic capabilities. However,
conventional nongated imaging techniques typically suffer from low
image quality or inadequate spatio-temporal resolution and fidelity.
Patient-Adaptive Reconstruction and Acquisition in Dynamic Imaging with
Sensitivity Encoding (PARADISE) is a highly accelerated nongated
dynamic imaging method that enables artifact-free imaging with high
spatio-temporal resolutions by utilizing novel computational techniques
to optimize the imaging process. In addition to using parallel imaging,
the method gains acceleration from a physiologically driven
spatio-temporal support model; hence, it is doubly accelerated. The
support model is patient adaptive, i.e., its geometry depends on
dynamics of the imaged slice, e.g., subject's heart rate and heart
location within the slice. The proposed method is also doubly adaptive
as it adapts both the acquisition and reconstruction schemes. Based on
the theory of time-sequential sampling, the proposed framework
explicitly accounts for speed limitations of gradient encoding and
provides performance guarantees on achievable image quality. The
presented in-vivo results demonstrate the effectiveness and feasibility
of the PARADISE method for high-resolution nongated cardiac MRI during
short breath-hold. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract161.txt#


Phase contrast MRI with multidirectional velocity encoding requires
multiple acquisitions of the same k-space lines to encode the
underlying velocities, which can considerably lengthen the total scan
time. To reduce scan time, parallel imaging is often applied. In
dynamic phase contrast MRI using standard generalized autocalibrating
partially parallel acquisitions (GRAPPA), several central k-spaces for
autocalibration of the reconstruction (autocalibrating signal lines
(ACS)) are typically acquired, separately for each velocity direction
and each cardiac timeframe, for calculating the reconstruction weights.
To further accelerate data acquisition, we developed two methods, which
calculated weights with a substantially reduced number of ACSl lines.
The effects on image quality and flow quantification were compared to
fully sampled data, standard GRAPPA, and time-interleaved sampling
scheme in combination with generalized autocalibrating partially
parallel acquisitions (TGRAPPA). The results show that the two proposed
methods can clearly improve scan efficiency while maintaining image
quality and accuracy of measured flow or myocardial tissue velocities.
Compared to TGRAPPA, the proposed methods were more accurate in
evaluating flow velocity. In conclusion, the proposed reconstruction
strategies are promising for dynamic multidirectionally encoded
acquisitions and can easily be implemented using the standard GRAPPA
reconstruction algorithm. Magn Reson Med, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract162.txt#


A new approach to autocalibrating, coil-by-coil parallel imaging
reconstruction, is presented. It is a generalized reconstruction
framework based on self-consistency. The reconstruction problem is
formulated as an optimization that yields the most consistent solution
with the calibration and acquisition data. The approach is general and
can accurately reconstruct images from arbitrary /i k/-space sampling
patterns. The formulation can flexibly incorporate additional image
priors such as off-resonance correction and regularization terms that
appear in compressed sensing. Several iterative strategies to solve the
posed reconstruction problem in both image and /i k/-space domain are
presented. These are based on a projection over convex sets and
conjugate gradient algorithms. Phantom and in vivo studies demonstrate
efficient reconstructions from undersampled Cartesian and spiral
trajectories. Reconstructions that include off-resonance correction and
nonlinear l /sub 1/-wavelet regularization are also demonstrated. Magn
Reson Med, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract163.txt#


Recent improvements in parallel imaging have been driven by the use of
greater numbers of independent surface coils placed so as to minimize
aliasing along the phase-encode direction(s). However, gains from
increasing the number of coils diminish as coil coupling problems begin
to dominate and the ratio of acceleration gain to expense for multiple
receiver chains becomes prohibitive. In this work, we redesign the
spatial-encoding strategy in order to gain efficiency, achieving a
gradient encoding scheme that is complementary to the spatial encoding
provided by the receiver coils. This approach leads to "/i O/-space"
imaging, wherein the gradient shapes are tailored to an existing
surface coil array, making more efficient use of the spatial
information contained in the coil profiles. In its simplest form, for
each acquired echo the Z2 spherical harmonic is used to project the
object onto sets of concentric rings, while the X and Y gradients are
used to offset this projection within the imaging plane. The theory is
presented, an algorithm is introduced for image reconstruction, and
simulations reveal that /i O/-space encoding achieves high encoding
efficiency compared to sensitivity encoding (SENSE) radial k-space
trajectories, and parallel imaging technique with localized gradients
(PatLoc), suggesting that /i O/-space imaging holds great potential for
accelerated scanning. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract164.txt#


The inherent distortions in echo-planar imaging that arise due to
inhomogeneities in the static magnetic field can lead to difficulties
when attempting to obtain structurally accurate diffusion-tensor
imaging data. Parallel acceleration techniques can reduce the magnitude
of these distortions but do not remove them entirely. Images can be
corrected using a measured field map, but this is prone to error. One
approach to correcting for these distortions, referred to here as
"blip-reversed" echo-planar imaging, involves collecting a second set
of images with the phase encoding reversed. Here, a novel approach to
collecting blip-reversed echo-planar imaging data for diffusion-tensor
imaging is presented: a dual-echo sequence is used in which the
phase-encoding direction of the second echo is swapped compared to the
first echo. This allows benefits of the blip-reversed approach to be
exploited, with only a modest increase in scan time and, due to the
extra data acquired, no significant loss of signal-to-noise efficiency.
A novel approach to recombining blip-reversed data is also presented,
which involves refining the measured field map, using an algorithm to
minimize the difference between the corrected images. The field map
refinement is also applicable to conventionally acquired blip-reversed
sequences. Magn Reson Med, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract165.txt#


We present ultra high speed optical coherence tomography (OCT) with
multi-megahertz line rates and investigate the achievable image
quality. The presented system is a swept source OCT setup using a
Fourier domain mode locked (FDML) laser. Three different FDML-based
swept laser sources with sweep rates of 1, 2.6 and 5.2MHz are compared.
Imaging with 4 spots in parallel quadruples the effective speed,
enabling depth scan rates as high as 20.8 million lines per second.
Each setup provides at least 98dB sensitivity and ~10 mu m resolution
in tissue. High quality 2D and 3D imaging of biological samples is
demonstrated at full scan speed. A discussion about how to best specify
OCT imaging speed is included. The connection between voxel rate, line
rate, frame rate and hardware performance of the OCT setup such as
sample rate, analog bandwidth, coherence length, acquisition dead-time
and scanner duty cycle is provided. Finally, suitable averaging
protocols to further increase image quality are discussed.

#e1#

#s1#
#abstract166.txt#


We have grown InN on 4H-SiC (0001) substrates with various off-angles
by RF-N/sub 2/ plasma molecular beam epitaxy (RF-MBE). Scanning
electron microscope observation revealed that InN films grown on 4H-SiC
(0001) substrates with off-angles of 4 degrees and 8 degrees are very
smooth and that there are no voids which have often observed for InN
epitaxial layers. X-ray diffraction reciprocal space maps for InN grown
on 4H-SiC (0001) showed that the c-axes of InN grown on 4H-SiC 4
degrees and 8 degrees off substrates are inclined by 0.35 degrees and
0.8 degrees , respectively, toward the misorientation of the substrate
while the c-axis of InN is parallel to that of 4H-SiC for the on-axis
substrate. Strong PL peak was observed from InN grown on 4 degrees off
substrate at 0.68 eV at 15 K. The PL peak was clearly observed even at
room temperature and simply shifted to lower energies with increasing
temperature. The difference in the PL peak energy between at 15 K and
300 K was 20 meV, which is reasonable taking into account the
difference in the thermal coefficients of InN and SiC. c WILEY-VCH
Verlag GmbH & Co. KGaA, Weinheim.

#e1#

#s1#
#abstract167.txt#


One-dimensional nanostructures such as silicon nanowires (SiNW) are
attractive candidates for low power density electronic and
optoelectronic devices including sensors. A new simple method for SiNW
bulk synthesis [1,2] is demonstrated in this work, which is inexpensive
and uses low toxicity materials, thereby offering a safe, energy
efficient and green approach. The method uses low flammability liquid
phenylsilanes, offering a safer avenue for SiNW growth compared with
using silane gas. A novel, duo-chamber glass vessel is used to create a
low-pressure environment where SiNWs are grown through
vapor-liquid-solid mechanism using gold nanoparticles as a catalyst.
The catalyst decomposes silicon precursor vapors of diphenylsilane and
triphenylsilane and precipitates single crystal SiNWs, which appear to
grow parallel to the substrate surface. This opens up possibilities for
synthesizing nano-junctions amongst wires which is important for the
grid architecture of nanoelectronics proposed by Likharev [3]. Even
bulk synthesis of SiNW is feasible using sacrificial substrates such as
Ca

#e1#

#s1#
#abstract168.txt#


Active Storage is a technology aimed at reducing the bandwidth
requirements of current supercomputing systems, and leveraging the
processing power of the storage nodes used by some modern file systems.
To achieve both objectives, Active Storage moves certain processing
tasks to the storage nodes, near the data they manage. Our proposal for
Active Storage has several key features: user-space implementation
which facilitates the port to different file systems, analytical model
to anticipate the performance of Active Storage with respect to a
traditional system, support for striped files and complex-format files
such as netCDF, and scientific-friendly programming and run-time
environment. [All rights reserved Elsevier].

#e1#

#s1#
#abstract169.txt#


I studied novel performance metrics of parallel processing efficiency,
determined simply by measuring timing data in a parallel process. The
metrics, which consider load-imbalance, are called the parallel
efficiency, load-balancing, and parallel impediment metrics. The
relationships between these three can be unified into one expression
that makes it possible to conduct detailed performance evaluations of
parallel processing. The parallel efficiency metric also makes it
possible to estimate the acceleration potential of the parallel
process. [All rights reserved Elsevier].

#e1#

#s1#
#abstract170.txt#


This paper presents a parallel differential correlation acquisition
algorithm in time domain for the GNSS hardware receivers, which can
acquire multi-satellite signals simultaneously by reusing the
correlation results. Both the computation complexity and the number of
registers needed by the proposed algorithm, the conventional
correlation algorithm, two modified correlation algorithms in time
domain and the FFT-based correlation algorithm in frequency domain are
also analyzed. Analysis and simulation results indicate that the
computation complexity of the proposed algorithm does not increase
while oversampling rate increases and the proposed algorithm has
advantages in computation complexity and is feasible in practical
hardware systems.

#e1#

#s1#
#abstract171.txt#


This paper considers single-stage multi-parallel machines scheduling
problems. A fuzzy scheduling model is developed with objective to
minimize the makespan. In the scheduling model, the uncertainty of
processing time is represented by triangular fuzzy number, and the
credibility theory is used to assess fuzzy number. In the optimal
scheduling searching process, we introduce matching theory to model the
scheduling problem consider in this paper, and two heuristic rules are
used to improve searching efficiency. At last, genetic algorithm is
developed for solving the scheduling problem, and numerical example
shows good result.

#e1#

#s1#
#abstract173.txt#


Dual comb-type electrodes were developed as a plasma source in very
high frequency (VHF) plasma enhanced chemical vapor deposition system
for uniform deposition of silicon films. Two VHF powers introduced to
each electrode produced parallel plasma bands, and their positions
could be changed by manipulating the phase difference between the
supplied VHF waves. Excitation frequency was 80 MHz. The maximum plasma
density using this plasma source was 1.5*10/sup 10//cm/sup 3/ and the
electron temperature was around 2 eV with input power of 2.5 kW, which
were measured by double tip Langmuir probe. The uniformity of
deposition rate under +/-13% was achieved on 1 m/sup 2/ area with
optimal plasma conditions. [All rights reserved Elsevier].

#e1#

#s1#
#abstract174.txt#


In this article we adopt the graphical processing units as a low-cost
and efficient solution of challenging electromagnetic numerical
problems. Based on the compute unified device architecture, an
optimized method of moments algorithm has been implemented which adopts
direct solvers based on LU decomposition. Numerical results obtained on
the practical case of a patch antenna analysis demonstrate the high
performance of the approach. c Wiley Periodicals, Inc.

#e1#

#s1#
#abstract175.txt#


In this article, a triplexer is presented for three operation
frequencies in GSM, based on stepped impedance resonators (SIR).SIR
structures have some characteristics, such as low weight, little
dimension, and the ability of providing filter behavior with different
frequencies in same dimensions. To design the triplexer, bandpass
filters of each output should be design first. Parallel connection of
three filters together causes an extreme increase in insertion loss of
each filter and also reflection of power to the input. A novel
based-on-SIR matching circuit is proposed to solve these problems. By
connecting optimized matching circuits to three bandpass filters [1],
which are simulated for three frequency ranges in GSM, the desired
triplexer is obtained. This based-on-SIR triplexer is simulated and
fabricated. The measurement results of the fabricated prototype shows a
good agreement with simulation results. c Wiley Periodicals, Inc.

#e1#

#s1#
#abstract176.txt#


Large scales of XML information comes continually from new Web
applications, and SLCA (Smallest Lowest Common Ancestor)-based XML
keyword search is one of the most important information retrieval
approaches. Previous approaches focus on building index for XML
documents. However in information dissemination scenario, it is
impossible to build index in advance for continuous XML document
streams. This paper addresses SLCA-based keyword search for continuous
XML documents by Map-Reduce mechanism. We use parallel algorithms to
process plenty of XML documents in Hadoop environment. A distributed
SLCA computation method is designed, where each net node computes SLCA
independently and just a little information needs be transmitted. A
real Hadoop environment is built and we demonstrate the efficiency of
our algorithms analytically and experimentally.

#e1#

#s1#
#abstract177.txt#


We investigate the problems of scheduling n weighted jobs to m
identical machines with availability constraints. We consider two
different models of availability constraints: the preventive model
where the unavailability is due to preventive machine maintenance, and
the fixed job model where the unavailability is due to a priori
assignment of some of the n jobs to certain machines at certain times.
Both models have applications such as turnaround scheduling or overlay
computing. In both models, the objective is to minimize the total
weighted completion time. We assume that m is a constant, and the jobs
are non-resumable. For the preventive model, it has been shown that
there is no approximation algorithm if all machines have unavailable
intervals even when w iota = pi for all jobs. In this paper, we assume
there is one machine permanently available and the processing time of
each job is equal to its weight for all jobs. We develop the first PTAS
when there are constant number of unavailable intervals. One main
feature of our algorithm is that the classification of large and small
jobs is with respect to each individual interval, thus not fixed. This
classification allows us (1) to enumerate the assignments of large jobs
efficiently; (2) and to move small jobs around without increasing the
objective value too much, and thus derive our PTAS. Then we show that
there is no FPTAS in this case unless P = NP. For fixed job model, we
first show that if job weights are arbitrary then there is no constant
approximation for a single machine with 2 fixed jobs or for two
machines with one fixed job on each machine, unless P = NP. As the
preventive model, we assume that the weight of a job is the same as its
processing time for all jobs. We show that the PTAS for the preventive
model can be extended to solve this problem when the number of fixed
jobs and the number of machines are both constants.

#e1#

#s1#
#abstract178.txt#


Orc language is a concurrency calculus proposed to study the
orchestration patterns in wide area computing. Its special properties
such as high concurrency and asynchronism makes it a brilliant subject
to study the distributed service oriented systems. This paper proposes
a denotational semantical model for Orc language. Every Orc program is
formalized to a predicate. Healthiness conditions are provided to make
the program domain corresponding to a specific subset of predicate
domain. This model gives the same semantical interpretation to the
implementations and specifications. With the refinement principle, we
are able to determine whether a program satisfies its specification,
which can be illustrated by theorem provers.

#e1#

#s1#
#abstract179.txt#


Machine readable open public data and the issue of multilingual web are
open challenges promising to transform the relationship between
citizens and European institutions. In this context the DALOS project
aims at ensuring coherence and alignment in the legislative language,
providing law-makers with knowledge management tools to improve the
control over the multilingual complexity of European legislation and
over the linguistic and conceptual issues involved in its transposition
into national laws. This paper describes the design and implementation
activities performed on the basis of a set of parallel texts in
different languages on a specific legal topic. Natural language
processing techniques have been applied to automatically build lexicons
for each language. Lexical and conceptual multilingual alignment has
been accomplished exploiting terms position in parallel documents. An
ontology describing entities involved in the chosen domain has been
developed in order to provide a semantic description of terms in
lexicons. A modular integration of such resources, represented in
RDF/OWL standard format, allowed their effective and flexible access
from a legislative drafting application prototype, able to enrich legal
documents with terms mark-up and semantic annotations.

#e1#

#s1#
#abstract180.txt#


Depth estimation in a scene using image pairs acquired by a stereo
camera setup, is one of the important tasks of stereo vision systems.
The disparity between the stereo images allows for 3D information
acquisition which is indispensable in many machine vision applications.
Practical stereo vision systems involve wide ranges of disparity
levels. Considering that disparity map extraction of an image is a
computationally demanding task, practical real-time FPGA based
algorithms require increased device utilization resource usage,
depending on the disparity levels operational range, which leads to
significant power consumption. In this paper a new hardware-efficient
real-time disparity map computation module is developed. The module
constantly estimates the precisely required range of disparity levels
upon a given stereo image set, maintaining this range as low as
possible by verging the stereo setup cameras axes. This enables a
parallel-pipelined design, for the overall module, realized on a single
FPGA device of the Altera Stratix IV family. Accurate disparity maps
are computed at a rate of more than 320 frames per second, for a stereo
image pair of 640*480 pixels spatial resolution with a disparity range
of 80 pixels. The presented technique provides very good processing
speed at the expense of accuracy, with very good scalability in terms
of disparity levels. The proposed method enables a suitable module
delivering high performance in real-time stereo vision applications,
where space and power are significant concerns. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract181.txt#


Discrete Cosine Transform (DCT) plays an important role in the image
and video compression, and it has been widely used in JPEG, MPEG,
H.26x. DCT being implemented by hardware is crucial to improve the
speed of image compression. This paper presents a method that 2-D DCT
is implemented by FPGA, which is based on the algorithm of row-column
decomposition, and the parallel structure is used to achieve high
throughput. The design is achieved by top-down design methodology and
described with Verilog HDL in RTL level. The hardware of 2-D DCT is
implemented by the FPGA EP2C35F672C8 made by ALTERA. The experiment
results show that the delay time is as low as 15 ns, and the clock
frequency as high as 138.35 MHz, which can satisfy the requirements of
the real-time video image compression.

#e1#

#s1#
#abstract182.txt#


Purpose - Material formulation, structuring and modification are key to
increasing the unit volume complexity and density of next generation
electronic packaging products. Laser processing is finding an
increasing number of applications in the fabrication of these advanced
microelectronic devices. The purpose of this paper is to discuss the
development of new laser-processing capabilities involving the
synthesis and optimization of materials for tunable device
applications. Design/methodology/approach - The paper focuses on the
application of laser processing to two specific material areas, namely
thin films and nanocomposite films. The examples include BaTiO/sub
3/-based thin films and BaTiO/sub 3/ polymer-based nanocomposites.
Findings - A variety of new regular and random 3D surface patterns are
highlighted. A frequency-tripled Nd:YAG laser operating at a wavelength
of 355 nm is used for the micromachining study. The micromachining is
used to make various patterned surface morphologies. Depending on the
laser fluence used, one can form a "wavy," random 3D structure, or an
array of regular 3D patterns. Furthermore, the laser was used to
generate free-standing nano and micro particles from thin film
surfaces. In the case of BaTiO/sub 3/ polymer-based nanocomposites,
micromachining is used to generate arrays of variable-thickness
capacitors. The resultant thickness of the capacitors depends on the
number of laser pulses applied. Micromachining is also used to make
long, deep, multiple channels in capacitance layers. When these
channels are filled with metal, the spacings between two metallized
channels acted as individual vertical capacitors, and parallel
connection eventually produce vertical multilayer capacitors. For a
given volume of capacitor material, theoretical capacitance
calculations are made for variable channel widths and spacings. For
comparison, calculations are also made for a "normal" capacitor, that
is, a horizontal capacitor having a single pair of electrodes. Research
limitations/implications - This technique can be used to prepare
capacitors of various thicknesses from the same capacitance layer, and
ultimately can produce variable capacitance density, or a library of
capacitors. The process is also capable of making vertical 3D
multilayer embedded capacitors from a single capacitance layer. The
capacitance benefit of the vertical multilayer capacitors is more
pronounced for thicker capacitance layers. The application of a laser
processing approach can greatly enhance the utility and optimization of
new materials and the devices formed from them. Originality/value -
Laser micromaching technology is developed to fabricate several new
structures. It is possible to synthesize nano and micro particles from
thin film surfaces. Laser micromachining can produce a variety of
random, as well as regular, 3D patterns. As the demand grows for
complex multifunctional embedded components for advanced organic
packaging, laser micromachining will continue to provide unique
opportunities.

#e1#

#s1#
#abstract183.txt#


Often server systems do not implement the best known algorithms for
optimizing average Quality of Service (QoS) out of concern that these
algorithms may be insufficiently fair to individual jobs. The standard
method for balancing average QoS and fairness is to optimize the L/sub
p/ norm, 1 # p # infinity . Thus we consider server scheduling
strategies to optimize the L/sub p/ norms of the standard QoS measures,
flow and stretch. We first show that there is no n degrees /sup
(1)/-competitive online algorithm for the L/sub p/ norms of either flow
or stretch. We then show that the standard clairvoyant algorithms for
optimizing average QoS, Shortest Job First (SJF), and Shortest
Remaining Processing Time (SRPT), are scalable for the L/sub p/ norms
of flow and stretch. We then show that the standard nonclairvoyant
algorithm for optimizing average QoS, Shortest Elapsed Time First
(SETF), is also scalable for the L/sub p/ norms of flow. We then show
that the online algorithm, Highest Density First (HDF), and the
nonclairvoyant algorithm, Weighted Shortest Elapsed Time First (WSETF),
are scalable for the weighted L/sub p/ norms of flow. These results
suggest that the concern that these standard algorithms may
unnecessarily starve jobs is unfounded. In contrast, we show that the
Round Robin, or Processor Sharing, algorithm, which is sometimes
adopted because of its seeming fairness properties, is not O(1 +
e)-speed, n degrees /sup (1)/-competitive for sufficiently small isin .

#e1#

#s1#
#abstract184.txt#


This paper presents an automatic image-based modelling method based on
shape from silhouettes that does not need any user interactions of
camera calibration or image segmentation. Under circular motion
constraints, using an iterative optimisation of graph cuts and
conjugate direction minimisation, we can label an object's visual hull
and minimise silhouette coherence, which is the difference between
projected visual hull and the background-subtracted silhouettes. This
process can converge to accurate camera parameters and smooth visual
hull. Using Graphics Processing Unit and its parallel computation
ability, our approach is not only automatic but also efficient, and can
produce realistic 3D models.

#e1#

#s1#
#abstract185.txt#


The PUPIL system is a combination of software and protocols for the
systematic linkage and interoperation of molecular dynamics and quantum
mechanics codes to perform QM/MD (sometimes called QM/MM) calculations.
The Gaussian03 and Amber packages were added to the PUPIL suite
recently. However, efficient parallel QM codes are critical because
calculation of the QM forces is the overwhelming majority of the
computational load. Here we report details of incorporation of the
deMon2k density functional suite as a new parallel QM code. An
additional motivation is to add a highly optimized, purely DFT code. We
illustrate with a demonstration study of the influence of perchlorate
as a dopant ion of the poly(3,4-ethylenedioxythiophene) conducting
polymer in explicit acetonitrile solvent using Amber and deMon2k. We
discuss unanticipated requirements for use of a scheme for
semi-empirical correction of Kohn-Sham eigenvalues to give physically
meaningful one-electron gap energies. We provide comparison of both
geometric parameters and electronic properties for nondoped and doped
systems. We also present results comparing deMon2k and Gaussian03
calculation of forces for a short sequence of steps. We discuss briefly
some difficult problems of quantum zone SCF convergence for the
anionically doped system. The difficulties seem to be caused by
well-know deficiencies in simple approximate exchange-correlation
functionals. c Wiley Periodicals, Inc.

#e1#

#s1#
#abstract186.txt#


We describe an implementation of a parallel document clustering scheme
based on latent semantic indexing, which uses singular value
decomposition. Given a set of documents, the clustering algorithm is
dynamic in the sense that it automatically infers the number of
clusters to be output. The parallel version has been implemented on a
LAN and on a dual-core system. Experimental evaluation of the algorithm
shows an average speed-up of 6.22 for the LAN implementation and an
average speed-up of 3.71 for the dual-core implementation, while still
maintaining a precision and recall in the range [0.85, 1]. To put these
implementations in the context of information retrieval, we use the
parallel clustering algorithm and develop a document similarity search
system. The similarity search system shows good performance in terms of
precision and recall. Copyright c John Wiley & Sons, Ltd.

#e1#

#s1#
#abstract188.txt#


A transposition graph is a Cayley graph in which each vertex
corresponds to a permutation and an edge is placed between permutations
if they differ by exactly one transposition. In this article, we
propose an efficient algorithm to find a collection of vertex-disjoint
paths connecting a given source vertex /i s/ and a given set of
destination vertices /i D/. The running time of the algorithm is
polynomial in the number of destination vertices, and the resultant
path connecting /i s/ and each destination is longer than the distance
to the destination by at most 16. c 2009 Wiley Periodicals, Inc.

#e1#

#s1#
#abstract189.txt#


We describe an efficient, high-level abstraction, multi-port
memory-control unit (MCU) capable of providing data at maximum
throughput. This MCU has been developed to take full advantage of FPGA
parallelism. Multiple parallel processing entities are possible in
modern FPGA devices, but this parallelism is lost when they try to
access external memories. To address the problem of multiple entities
accessing shared data we propose an architecture with multiple abstract
access ports (AAPs) to access one external memory. Bearing in mind that
hardware designs in FPGA technology are generally slower than memory
chips, it is feasible to build a memory access scheduler by using a
suitable arbitration scheme based on a fast memory controller with AAPs
running at slower frequencies. In this way, multiple processing units
connected through the AAPs can make memory transactions at their slower
frequencies and the memory access scheduler can serve all these
transactions at the same time by taking full advantage of the memory
bandwidth. [All rights reserved Elsevier].

#e1#

#s1#
#abstract190.txt#


Parallel control and management have been proposed as a new mechanism
for conducting operations of complex systems, especially those that
involved complexity issues of both engineering and social dimensions,
such as transportation systems. This paper presents an overview of the
background, concepts, basic methods, major issues, and current
applications of Parallel transportation Management Systems (PtMS). In
essence, parallel control and management is a data-driven approach for
modeling, analysis, and decision-making that considers both the
engineering and social complexity in its processes. The developments
and applications described here clearly indicate that PtMS is effective
for use in networked complex traffic systems and is closely related to
emerging technologies in cloud computing, social computing, and
cyberphysical-social systems. A description of PtMS system
architectures, processes, and components, including OTSt, Dyna CAS,
aDAPTS, iTOP, and TransWorld is presented and discussed. Finally, the
experiments and examples of real-world applications are illustrated and
analyzed.

#e1#

#s1#
#abstract191.txt#


One of the main problems of substructure-based parallel solution
methods is the imbalances in the condensation times of substructures
when direct solvers are used. Such imbalances usually decrease the
performance of the parallel solution. Thus, in this study, a workload
distribution framework for such methods at heterogeneous computing
environment is presented. The main idea behind this framework is to
iteratively adjust the shapes of substructures so that the imbalance in
their condensation times is minimized. Both generated and actual
structural models were solved to illustrate the applicability and the
efficiency of this approach. In these examples, a PC cluster having
eight different computers was used.

#e1#

#s1#
#abstract192.txt#


The paper presents a workflow application aimed at detection of
unwanted and potentially dangerous events observed by multimedia
cameras installed at public facilities, such as universities. The
detection of unwanted events launches alarms, causes notifications to
be sent to respective forces and videos to be recorded. The proposed
application model consists of several stages: nodes with camera(s)
attached to them, node(s) for fetching data streams, node(s) for data
processing and nodes used for notification and data storage. The
already implemented testbed application fetches either a stream of data
from a camera attached to a computer or reads a video file. The stream
is decoded and data are processed by parallel processes on a cluster.
The implemented testbed application detects a movement in the stream
and can be applied for monitoring the university during the night.

#e1#

#s1#
#abstract193.txt#


A matrix times vector multiplication (matvec) is a cornerstone
operation in iterative methods of solving large sparse systems of
equations such as the conjugate gradients method (cg), the minimal
residual method (minres), the generalized residual method (gmres) and
exerts an influence on overall performance of those methods. An
implementation of matvec is particularly demanding when one executes
computations on a GPU (Graphics Processing Unit), because using this
device one has to comply with certain programming rules in order to
take advantage of parallel computing. In this paper, it will be shown
how to modify the sparse matrix-vector multiplication based on CRS
(Compressed Row Storage) to achieve about 3-5 times better performance
on - a low cost - GPU (GeForce GTX 285, 1.48 GHz) than on a CPU (Intel
Core i7, 2.67 GHz).

#e1#

#s1#
#abstract194.txt#


In pervasive computing paradigm, the image data obtained by a variety
of multimedia, information equipment. In order to make use of these
information reasonably and efficiently, a new method for image fusion
based on fuzzy neural network has presented in this paper. It can fuse
massive data from the multi-sensor image. Here, fuzzy neural network is
a parallel information processing model, there are more adaptive and
self-organization, and can accomplish the complexity of real-time
computing and mass data retrieval, it demonstrate its unique
superiority of image understanding, pattern recognition and the
handling of incomplete information. The fuzzy neural network system
developed by us can be used in multi-sensor image fusion. Based on our
experiments, it has been proved that the fusion is fast, effective, and
can meet the real-time requirements of pervasive computing.

#e1#

#s1#
#abstract195.txt#


With the proliferation of internet technologies, publish/subscribe
systems have gained wide usage as a middleware. However for this model,
catering large number of publishers and subscribers while retaining
acceptable performance is still a challenge. Therefore, this paper
presents two parallelization strategies to improve message delivery of
such systems. Furthermore, we discuss other techniques which can be
adopted to increase the performance of the middleware. Finally, we
conclude with an empirical study, which establishes the comparative
merit of those two parallelization strategies in contrast to serial
implementations.

#e1#

#s1#
#abstract196.txt#


The monitoring and control of environmental conditions in fabrication
areas has become an important issue due to the newest technologies that
are usually very complex and very sensitive to the external influences
produced by humidity, dust particles, temperature and pressure
variations etc. In this context, this paper presents a data acquisition
system that is capable to monitor and measure three types of
environmental parameters: pressure, temperature, and humidity. The
remote board containing sensors and processing circuits is connected to
the PC through parallel port. The analog to digital conversion is
realized with a resolution of 8 bits using the ADC0804 circuit. The
systems can operate in two specific modes: monitoring and measurement.
The sensors are read periodically with a selectable frequency and the
selection of the sensors is realized with three signals of parallel
port: STROBE for temperature, AUTOFD for pressure and INIT for
humidity. The board control and the acquired data processing are
realized with a LabVIEW application that is capable to simultaneously
display the measured data. With minimal changes, the proposed system
can be extended to operate with more types of sensors. Using a wireless
transmission method between PC and remote board, the operation distance
of the system can be further extended.

#e1#

#s1#
#abstract197.txt#


The goal of this paper is to investigate a new method of parallel
computing used for a wind speed model, taking advantage of a modern,
multicore CPU and of LabVIEW graphical programming language. The wind
speed model is widely used in literature as a superposition of two
components, a low and a high frequency component. The high frequency or
turbulent component is usually generated by filtering a white noise
sequence with a shaping filter. The shaping filter is based on von
Karman's model and is of a non-integer order, so it requires a
computational effort. The wind speed model is first generated using
Matlab. A multicore application is proposed, which uses the pipelining
technique and all 4 cores of the processor. A comparison between the
two methods is made, showing the performance gain of the
parallelization, compared to a traditional, sequential computing
approach.

#e1#

#s1#
#abstract198.txt#


In this paper the idea of the general-purpose processor implemented in
dynamically reconfigurable FPGA is presented. The novelty of the
proposed solution lays in the lack of typical sequential processing -
all operations are realized in parallel in the hardware. At the same
time the new architecture does not impose any modification of the
software development process.

#e1#

#s1#
#abstract199.txt#


This paper presents a digital, transistor level implemented neo-fuzzy
neural network. This type of neural network is particularly well suited
for real-time applications like those encountered in signal processing
and nonlinear system identification. We consider in detail a flexible
reconfigurable circuit of a single nonlinear synapse of this network.
When combining such circuits, single-layer or multilayer networks can
be designed. The advantages of the proposed circuit come in the form of
reduced redundancy, high data rate due to parallel operation, low power
consumption, and an overall flexibility of system configuration.

#e1#

#s1#
#abstract200.txt#


This paper presents the design of a digital CMOS integrated circuit
implementing a type-2 fuzzy logic controller. The proposed architecture
is suitable for serial processing of fuzzy rules combined with parallel
processing of upper and lower membership functions of type-2 fuzzy
sets. The parameterized VHDL model allows to synthesize the circuit of
the required size for a particular application. Moreover, on-chip
programming is performed.

#e1#

#s1#
#abstract201.txt#


Digital architecture of fuzzy processor is proposed. All blocks - fuzzy
sets (triangular), rule strength calculation (minimum) and
defuzzyfication (weighted sum) were implemented in VHDL, verified and
synthesized for FPGA. Implementation of floating point division block
appeared to be the most difficult part of the design. Partially
concurrent and pipelined data flow provides competitive performance,
with relatively little dependence on particular algorithm complexity.

#e1#

#s1#
#abstract202.txt#


Neuromorphic circuits try to replicate aspects of the information
processing in neural tissue. Historically, this has often meant some
kind of long-term learning function which slowly adjusts the weight of
a synapse to achieve a certain target network function. Recently,
short-term dynamics at the synapse have also gained significant
attention due to their role in dynamic and temporal information
processing. However, only very few neuromorphic circuits have
incorporated short term dynamics, with still fewer of these
implementations being biologically realistic. We derive a circuit for
biologically relevant short term dynamics, showing its accuracy with
respect to biological measurements. Since this circuit significantly
increases the overall complexity of the synapse, a direct integration
in the synapse would be prohibitive. Thus, in addition to the short
term dynamics, we also present a novel configurable topology for the
neurons and synapses on chip which achieves a compact and flexible
overall design while still augmenting all synapses with the new short
term dynamics.

#e1#

#s1#
#abstract203.txt#


This paper describes the development of a FPGA-based object detection
algorithm for manipulation purposes in a mobile robot. The target
application is a robotic system which aids workers in a manufacturing
plant. The whole system is provided with a camera which captures images
of the objects that can be found in the environment. The FPGA extracts
the most useful data from these images and performs the object
recognition tasks by means of a neural network. For performance
reasons, the neural network is implemented in the hardware partition of
the system, while the rest of the algorithms is included in an embedded
processor. This design provides a tradeoff between the flexibility and
accuracy of the software in performing image processing algorithms and
the high-speed of the hardware to execute parallel computations, useful
for the neural network.

#e1#

#s1#
#abstract204.txt#


We describe a CMOS image sensor with column-parallel delta-sigma (
Delta Sigma ) analog-to-digital converter (ADC). The design employs
three transistor pixels (3T/sup 1/) where the unique configuration of
the Delta Sigma ADC reduces the noise contribution of the readout
transistor. A 128 * 128 pixel image sensor prototype is fabricated in
0.35 mu m TSMC technology. The reset noise and the offset fixed pattern
noise (FPN) are removed in the digital domain. The measured readout
noise is 37.8 mu V for an exposure time of 33 ms. The low readout noise
allows an improved low light response in comparison to other
state-of-art designs. The design is suitable for applications demanding
excellent low-light response such as astronomical imaging. The sensor
has a measured intra-scene dynamic range (DR) of 91 dB, and a peak
signal-to-noise ratio (

#e1#

#s1#
#abstract205.txt#


In this letter, a novel pipelined decision feedback RNN equalizer
(PDFRNE) with low computational complexity is proposed. Since each
module is a DFRNN with the decision feedback structure so that it can
eliminate the past error remaining in the network. Moreover, the
performance can be further improved. At the same time, it can overcome
the unstableness due to its nature of the infinite impulse response
(IIR) structure.

#e1#

#s1#
#abstract206.txt#


The work aims at the experimental and theoretical study of the
mechanism of meltblowing. Meltblowing is a popular method of producing
polymer microfibers and nanofibers en masse in the form of nonwovens
via aerodynamic blowing of polymer melt jets. However, its physical
aspects are still not fully understood. The process involves a complex
interplay of the aerodynamics of turbulent gas jets with strong
elongational flows of polymer melts, none of them fully uncovered and
explained. To evaluate the role of turbulent pulsations (produced by
turbulent eddies in the gas jet) in meltblowing, we studied first a
model experimental situation where solid flexible sewing threadlines
were subjected to parallel high speed gas jet. After that a
comprehensive theory of meltblowing is developed, which encompasses the
effects of the distributed drag and lift forces, as well as turbulent
pulsations acting on polymer jets, which undergo, as a result, severe
bending instability leading to strong stretching and thinning.
Linearized theory of bending perturbation propagation over threadlines
and polymer jets in meltblowing is given and some successful
comparisons with the experimental data are demonstrated.

#e1#

#s1#
#abstract207.txt#


We investigate a fast pedestrian localization framework that integrates
the cascade-of-rejectors approach with the Histograms of Oriented
Gradients (HoG) features on a data parallel architecture. The salient
features of humans are captured by HoG blocks of variable sizes and
locations which are chosen by the AdaBoost algorithm from a large set
of possible blocks. We use the integral image representation for
histogram computation and a rejection cascade in a sliding-windows
manner, both of which can be implemented in a data parallel fashion.
Utilizing the NVIDIA CUDA framework to realize this method on a
Graphics Processing Unit (GPU), we report a speed up by a factor of 13
over our CPU implementation. For a 1280*960 image our parallel
technique attains a processing speed of 2.5 to 8 frames per second
depending on the image scanning density, which is similar to the recent
GPU implementation of the original HoG algorithm in.

#e1#

#s1#
#abstract208.txt#


We present an integral image algorithm that can run in real-time on a
Graphics Processing Unit (GPU). Our system exploits the parallelisms in
computation via the NIVIDA CUDA programming model, which is a software
platform for solving non-graphics problems in a massively parallel
high-performance fashion. This implementation makes use of the
work-efficient scan algorithm that is explicated in. Treating the rows
and the columns of the target image as independent input arrays for the
scan algorithm, our method manages to expose a second level of
parallelism in the problem. We compare the performance of the parallel
approach running on the GPU with the sequential CPU implementation
across a range of image sizes and report a speed up by a factor of 8
for a 4 megapixel input. We further investigate the impact of using
packed vector type data on the performance, as well as the effect of
double precision arithmetic on the GPU.

#e1#

#s1#
#abstract209.txt#


In this paper we will present a geometric lane representation which
enables a 4D-lane tracker to handle ambiguous, non-parallel and even
crossing lane marking tracks in complex situations. The introduced
geometry model contains lateral independent tracks which lie on top of
the surface of a continuous shape that is formed like the road surface
within a specific preview and is bent in horizontal, vertical and
torsional way. Due to a more precise shape representation the
association of measured to expected features is eased and therefore the
approach leads to a noticeable improvement of stability of the tracking
system. The introduced state description achieves a decoupling of
dynamic ego movement and stationary lane markings to enhance the
robustness furthermore. Based on the increased number of estimated
tracks the computing effort of the feature extraction rises linearly
and the estimation effort approximately quadratically. This challenge
is met by a selective innovation mechanism which allows a comparable
calculation time and is presented below as well.

#e1#

#s1#
#abstract210.txt#


A parallel implementation of the split-step Fourier method utilizing
the general purpose parallel computing architecture for graphics
processing units CUDA is presented. Results of the GPU-implementation
are compared to a conventional CPU-based approach regarding computation
time and accuracy. We developed a novel implementation with a
significantly higher accuracy than the CUDA intrinsic FFT in single
precision mode yielding a high speed-up factor of up to 144 compared to
a CPU implementation.

#e1#

#s1#
#abstract211.txt#


Some initial investigations are conducted to apply Ant Colony Algorithm
(A

#e1#

#s1#
#abstract212.txt#


Rapid changes on the goods market take place today. These changes
require an increased flexibility of the production systems in order to
quickly adapt the manufacturing system to a new model of product.
However, increasing the flexibility usually also increases the
complexity of the system, and therefore the costs per part produced.
The challenge is to develop control strategies capable to increase the
productivity of a production system, while maintaining the associated
costs as low as possible. We propose a control strategy capable of
responding to this challenge. Using multitasking programming and
optimization of the materials' travel between different points of the
flexible manufacturing system, we implement a controller capable to
reduce the lead time by up to 26.3%, which leads to an increase of the
system's productivity by up to 34%. The control strategy proposed will
increase the productivity; moreover no interventions on the hardware
systems are required. The strategy also reduces the manufacturing cost
per part since the productivity increases.

#e1#

#s1#
#abstract213.txt#


We present a simple and versatile patterning procedure for the reliable
and reproducible fabrication of high aspect ratio (10 /sup 4/ )
electrical interconnects that have separation distances down to 20 nm
and lengths of several hundreds of microns. The process uses standard
optical lithography techniques and allows parallel processing of many
junctions, making it easily scalable and industrially relevant. We
demonstrate the suitability of these nanotrenches as electrical
interconnects for addressing micro and nanoparticles by realizing
several circuits with integrated species. Furthermore, low impedance
metal-metal low contacts are shown to be obtained when trapping a
single metal-coated microsphere in the gap, emphasizing the intrinsic
good electrical conductivity of the interconnects, even though a wet
process is used. Highly resistive magnetite-based nanoparticles
networks also demonstrate the advantage of the high aspect ratio of the
nanotrenches for providing access to electrical properties of highly
resistive materials, with leakage current levels below 1 pA.

#e1#

#s1#
#abstract214.txt#


Solakli Basin is located Eastern Black Sea region where high mountain
ranges run parallel to the coast in the north. Such a mountainous
terrain, it is generally hard or impossible to reach to acquire data by
terrestrial measurement. However, today, by using integration of Remote
Sensing and Geographic Information Systems even those kinds of basins
can be modelled. These techniques provide to derive basin, land use
and/or soil type characteristics in an accurate and quick way,
particularly for water resources assessment studies. In addition to
basin characteristics, spatial distribution of precipitation is also
important for these types of studies. In this study, for the
classification of Solakli Basin IRS P6 multispectral satellite data
with 5.8 m spatial resolution are used and to derive the Digital
Elevation Model IRS P5 stereo satellite data with 2.5 m spatial
resolution is used. The basin characteristics are mathematically
determined. Isohyetal maps to understand precipitation distribution are
generated by means of different geostatistical methods such as Inverse
Distance Weight, Radial Basis Function and Kriging. Among these
methods, Kriging and Radial Basis Function give more satisfactory
results.

#e1#

#s1#
#abstract1000.txt#


A high-throughput low-latency digital /i finite impulse response/ (FIR)
filter has been designed for use in /i partial-response
maximum-likelihood/ (PRML) read channels of modern disk drives. The
filter is a hybrid synchronous-asynchronous design. The speed-critical
portion of the filter is designed as a high-performance asynchronous
pipeline sandwiched between synchronous input and output portions,
making it possible for the entire filter to be embedded within a
clocked system. A novel feature of the filter is that the degree of
pipelining is dynamically variable, depending upon the input data rate.
This feature is critical in obtaining a very low filter latency
throughout the range of operating frequencies. The filter is a ten-tap
six-bit FIR filter, fabricated in a 0.18- mu m CMOS process. Resulting
chips were fully functional over a wide range of supply voltages, and
exhibited throughputs of over 1.3 giga-items/s, and latencies of 2-5
clock cycles. Interestingly, the filter throughput was limited by the
synchronous portion of the chip; the internal asynchronous pipeline was
estimated to be capable of significantly higher throughputs, around 1.8
giga-items/s. More importantly though, the adaptively pipelined nature
of the filter allows it to offer a worst-case latency of only 10 ns,
which is half the worst-case latency of the best previously reported
comparable fully-synchronous implementation by Rylov et al.

#e1#

#s1#
#abstract1001.txt#


Co/sub x/Zn/sub 1-x/ nanorod arrays were fabricated by
electrodeposition in porous anodic aluminum oxide templates at
different electric potentials. X-ray diffraction and transmission
electron microscopy indicate that highly-ordered and uniform nanorods
have been fabricated. The amounts of Co and Zn contents are
investigated using energy dispersive spectroscopy, which demonstrates
that the atom ratio of the alloy nanorods changes with the deposition
potential. In addition, magnetic measurements show that the magnetic
isotropy Co-rich CoZn nanorods will change to magnetic anisotropy
nanorods with the easy axis parallel to the rod long axis with
decreasing Co content.

#e1#

#s1#
#abstract1002.txt#


Although the similarities of synthetic aperture radar (SAR) and
computer aided tomography (CAT) imaging systems have been discussed in
several previously published papers, a rigorous implementation of
algorithms from one modality have not fully ventured into the other
modality(ies). This paper proposes image formation of synthetic
aperture radar (SAR) data using techniques developed from CAT systems.
The paper provides an overview of the typical signal processing
involved with SAR, B-mode ultrasound, and CAT imaging systems.
Simulations of SAR received echo data processed using techniques from
ultrasound and CAT imaging systems are provided to showcase the
possibilities. Further discussions of our proposed future work by
incorporating CAT techniques to inverse SAR and real time parallel
implementation using general purpose graphical processing units are
presented.

#e1#

#s1#
#abstract1003.txt#


In this paper we combine a new method of measuring infrared target
signature evolution with current research and developmental tracking
algorithms. Thermal images are decomposed by a set of Gabor filters and
demodulated to produce a set of spatio-spectrally localized AM-FM
functions corresponding to oriented texture regions from within the
original image. Critical updates are detected and issued to a particle
filter based tracker, operating only on a modulation domain target
model, by applying an empirically determined threshold to a new target
evolution measurement introduced in this paper. We achieve results
comparable to several other theoretical tracking algorithms at a
significantly reduced computational cost by eliminating the need to
perform parallel tracking in both the pixel and modulation domains.

#e1#

#s1#
#abstract1004.txt#


Real-time transient stability simulation is of paramount importance for
system security assessment and to initiate preventive control actions
before catastrophic events such as blackouts happen. Transient
stability simulation of realistic-size power systems involves the
solution of a large set of non-linear differential-algebraic equations
in the time-domain which requires significant computational resources.
Exploitation of parallel processing techniques can provide an efficient
and cost-effective solution to this problem. This paper proposes a
fully parallel method known as instantaneous relaxation (IR) for
real-time transient stability simulation. To validate and evaluate the
proposed method a test system has been implemented on a distributed
PC-Cluster based real-time simulator. A comparison of the captured
real-time results with those from the PSS/E software shows high
accuracy.

#e1#

#s1#
#abstract1005.txt#


In this paper we present a computing system which based on computations
on intervals over [0,1]. The interval-values are built up from points
and atomic intervals. The Boolean operators are extended to these
values in a natural way. Changing the bytes of the traditional
computing to interval-values one gets the interval-valued computing
device. We investigate the operators Rshift and Lshift as well. By a
simulation the interval-valued computing device can compute everything
that the classical computing devices can compute using bits in a byte.
Moreover, theoretically, this device is more effective than the
traditional computers, because of the possibility of infinite
parallelism. A method is presented to solve the Q-SAT problem in linear
time using the interval-valued approach. Our device uses the usual
concept of intervals to make the computations in a very effective way.
The list-representation of the interval-values is also presented. Based
on these lists, theoretically, one can simulate the
interval-computation both in mathematical and traditional computational
way.

#e1#

#s1#
#abstract1006.txt#


A real-time VHF swept frequency (20-300 MHz) reflectometry measurement
for radio-frequency capacitive-coupled atmospheric pressure plasmas is
described. The measurement is scalar, non-invasive and deployed on the
main power line of the plasma chamber. The purpose of this VHF signal
injection is to remotely interrogate in real-time the frequency
reflection properties of plasma. The information obtained is used for
remote monitoring of high-value atmospheric plasma processing.
Measurements are performed under varying gas feed (helium mixed with
0-2% oxygen) and power conditions (0-40 W) on two contrasting reactors.
The first is a classical parallel-plate chamber driven at 16 MHz with
well-defined electrical grounding but limited optical access and the
second is a cross-field plasma jet driven at 13.56 MHz with open
optical access but with poor electrical shielding of the driven
electrode. The electrical measurements are modelled using a lumped
element electrical circuit to provide an estimate of power dissipated
in the plasma as a function of gas and applied power. The performances
of both reactors are evaluated against each other. The scalar
measurements reveal that 0.1% oxygen admixture in helium plasma can be
detected. The equivalent electrical model indicates that the current
density between the parallel-plate reactor is of the order of 8-20 mA
cm/sup -2/. This value is in accord with 0.03 A cm/sup -2/ values
reported by Park et al (2001 J. Appl. Phys. 89 20-8). The current
density of the cross-field plasma jet electrodes is found to be 20
times higher. When the cross-field plasma jet unshielded electrode area
is factored into the current density estimation, the resultant current
density agrees with the parallel-plate reactor. This indicates that the
unshielded reactor radiates electromagnetic energy into free space and
so acts as a plasma antenna.

#e1#

#s1#
#abstract1022.txt#


In order to make MIT fast it is recommendable to excite several
transmit coils in parallel and to acquire all receive channels
simultaneously. The separation between the transmit channels can be
achieved e.g. by slightly separating the excitation frequencies, so
that it is possible to separate their contributions in the receivers by
synchronous demodulation. One major problem is the low output impedance
of the driver amplifiers so that each transmit coil acts as short
circuited when seen from the other transceivers and hence perturbs the
primary field. This causes a number of complications both in the
reconstruction software as well as from the viewpoint of

#e1#

#s1#
#abstract1023.txt#


Living cells react to external influences such as pharmacological
agents in an intricate manner due to their complex internal signal
processing. Cell reactions are an impact on vitality, cell-cell or
cell-matrix interaction and morphological changes. A number of
published techniques on impedance spectroscopy (IS) of adherent cells
with planar electrodes address these changes. However, IS can merely
serve as an indicator of cellular events rather than provide detailed
information on a specific cell process. Thus our approach is a
24-microwell sensor-plate with impedance-electrodes in parallel to pH-
and O/sub 2/-sensors, capable of being integrated into a fully
automated screening system. For the purpose of IS, high precision
impedance-electronics have been developed based on integrated circuits
and validated against a Solartron 1260 impedance analyzer. IS data is
correlated to the metabolic-sensors and additionally compared with cell
images shot by an inverse optical microscope which is also part of the
screening system. Proof of principle is demonstrated by experimental
growth monitoring of a MCF-7 culture and cellular response to
chemotherapeutics. Furthermore, the potential to monitor living tissue
probes is presented for the first time.

#e1#

#s1#
#abstract1024.txt#


We developed a new microscopic electrical impedance tomography
(micro-EIT) system to visualize admittivity distributions within a
miniature hexahedral container, where we place small biological samples
with a background solution or gel. Each of two facing sides (left and
right) of the container is fully covered by a solid metal electrode. We
inject current between them, thereby producing a uniform parallel
current flow along the longitudinal direction inside the container.
Each of three sides at the bottom, front and back is equipped with a
15*8 array of voltage-sensing electrodes. Switching modules are located
underneath the container so that we can measure voltage between any
neighbouring pair of electrodes. Three switching modules are connected
to a 16-channel multi-frequency EIT system to collect induced voltage
data from the three sets of 15x8 array electrodes subject to the single
fixed current injection. Voltage data set from 360 voltage-sensing
electrodes on three sides are utilized to produce cross-sectional
images of the admittivity distribution. We describe the design and
construction of the new micro-EIT system. Our future work should
include development of a customized image reconstruction algorithm for
the micro-EIT system and experimental validation.

#e1#

#s1#
#abstract1025.txt#


We adapt concepts from matched filtering to propose a method for
generating reconfigurable multiple beams. Combined with the Generalized
Phase Contrast (GPC) technique, the proposed method coined mGPC can
yield dynamically reconfigurable optical beam arrays with high light
efficiency for optical manipulation, high-speed sorting and other
parallel spatial light applications.

#e1#

#s1#
#abstract1026.txt#


The field of Quantum Imaging exploits the quantum nature of light and
the intrinsic parallelism of optical signals to devise novel techniques
for optical imaging and for parallel information processing at the
quantum level (see and references quoted therein). In this
presentation, we will shortly discuss two topics in this area: the
first is the so-called ghost imaging which, however, is not necessarily
quantum; the second is the detection of faint amplitude objects with a
sensitivity beyond the standard quantum limit, and in this case we are
fully in the quantum domain. Both topics are related to the phenomenon
of optical parametric down-conversion (PDC), in which a fraction of the
pump photons of a laser beam, injected into a crystal with a quadratic
non-linearity, are down-converted to a pair of signal and idler
photons, with conservation of total energy and total momentum. A
feature of paramount importance is that the signal and idler beams are
spatially correlated both in the near field (position correlation) and
in the far field (momentum correlation).The simultaneous presence of
position and momentum correlation implies quantum entanglement, as it
has been also observed experimentally.

#e1#

#s1#
#abstract1031.txt#


A continuum model of piezoelectric potential generated in a bent ZnO
nanorod cantilever is presented by means of the first piezoelectric
effect approximation. The analytical solution of the model shows that
the piezoelectric potential in the nanorod is proportional to the
lateral force but is independent along the longitudinal direction. The
electric potential in the tensile area and that in the compressive area
are antisymmetric in the cross section of the nanorod, which makes the
nanorod a `parallel plate capacitor' for piezoelectric nanodevices,
such as a nanogenerator. The magnitude of piezoelectric potential for a
ZnO nanorod of 50 nm diameter and 600 nm length bent by a 80 nN lateral
force is about 0.27 V, which is in good agreement with the finite
element method calculation.

#e1#

#s1#
#abstract1032.txt#


A human face does not only identify an individual but also communicates
useful information about a person's emotional state. No wonder
automatic face expression recognition has become an area of immense
interest within the computer science, psychology, medicine and
human-computer interaction research communities. Various feature
extraction techniques based on statistical to geometrical data have
been used for recognition of expressions from static images as well as
real time videos. In this paper we present a method for automatic
recognition of facial expressions from face images by providing
Discrete Wavelet Transform (DWT) features to a bank of five parallel
neural networks. Each neural network is trained to recognize a
particular facial expression, so that it is most sensitive to that
expression. Multi-classification is achieved by combining multiple
neural networks performing binary classification using oneagainst-all
approach. The outputs of all neural networks are combined using a
maximum function. The classification efficiency is tested on static
images from the publicly available JAFFE database. The experiments
using the proposed method demonstrate promising results.

#e1#

#s1#
#abstract1034.txt#


EMI engineers are struggling everyday with complex radiation problems
that fail critical products to pass EMI certification and causes big
loss of profit. Advances in EMI engineering are following a similar
trend like Signal-Integrity engineering 10-years ago when simulation
tools became capable of providing accurate predictive simulations in a
reasonable amount of time. With careful engineering utilizing
cutting-edge full-wave field-solver software: Momentum (MOM), EMpro
(FDTD) along with a hardware boost with heterogeneous massive CPU/GPU
parallel processing (CUDA) technology, we can move the EMI teams from
the back-end black-magic to a successful cost-effective front-end
design. This paper presents an innovative process (Virtual-EMI lab) for
pre- and post-tape-out providing the designers with an early stage
EMI-suppression matrix (on-chip and onboard enablers) to find the
optimum trade-off between performance and cost.

#e1#

#s1#
#abstract1035.txt#


Graphic Processing Units (GPUs) have evolved to provide a massive
computational power. In contrast to Central Processing Units, GPUs are
so-called many-core processors with hundreds of cores capable of
running thousands of threads in parallel. This parallel processing
power can accelerate the simulation of communication systems. In this
work, we utilize NVIDIA's Compute Unified Device Architecture (CUDA) to
execute two different sphere decoders on a graphic card. Both flat
fading and frequency selective channels are considered. We find that
the execution of the soft-sphere decoder can be accelerated by factors
of 6-8, and the fixed-complexity sphere decoder even by a factor of 50.

#e1#

#s1#
#abstract1036.txt#


This paper presents a simple and fast algorithm for labeling connected
components in binary images, based on a parallel label-broadcast
paradigm. A grid of processing units (called spiders) is used and each
element is responsible for updating its label value, during a specific
number of iterations. We describe the design and implementation of an
embedded architecture for real-time labeling of black and white images
based on FPGA technology. Since the image is divided and processed
independently by processing elements, it is possible to use the
proposed algorithm in an FPGA platform attached to an image sensor and
have a focal plane processor circuit-like.

#e1#

#s1#
#abstract1037.txt#


The FCT (Formatting, Communications and Temperatures) electronic unit
developed at INTA's radar laboratory as part of the X-Band synthetic
aperture radar project (RBX SAR, see ref) is capable of receiving up to
four channels of radar pulses and echoes at 1.5 Gbps line rates, with
the use of aurora protocol on top of the Virtex 5 gigabit transceivers,
format these data and send them to the system's storage unit (UAD) with
a parallel protocol. In addition, the FCT communicates with other units
through a VME bus. It also handles the memory map of the X-band
transmitting unit, as well as translates the VME parallel bus
communications to this unit to a 1.5 Gbps serial aurora protocol
through optical fiber. Besides, the FCT monitors and stores temperature
data gathered from different units of the radar system through a serial
SMBus protocol, and manages access to these data through VME bus.
Finally, with a second Virtex 5, it formats, decimates and filters the
received radar echoes and sends them to the real time processing unit
through a serial sFPDP protocol (1.0625 Gbaud).

#e1#

#s1#
#abstract1038.txt#


This paper presents the design and evaluation of architectures that
performs the SSD (Sum of Squared Differences) similarity criterion
calculation. The comparison was made with other widely used criterion:
the SAD (Sum of Absolute Differences). In order to compare the impact
of both criteria in the coding process, a set of executions using the
JM 16.0 reference software were performed. In these tests, SSD almost
ever got a better video quality than SAD. Three architectures are
proposed to perform SSD: (a) the first one uses a multiplexer, (b) the
second uses a memory and (c) the last one uses a dedicated multiplier.
One architecture to perform SAD is proposed to be compared with the
architectures using SSD. Each solution was described in VHDL and
synthesized to an Altera Stratix II FPGA. The video quality gain using
SSD over the SAD encourages the use of SSD calculators even with a
lower operation frequency when compared with an SAD implementation. In
the best case and considering HDTV 1080p videos (1920 * 1080 pixels),
it is possible to reach real time processing (30 frames per second) by
putting 12 SSD calculators working in parallel.

#e1#

#s1#
#abstract1039.txt#


This paper shows the effectiveness of a classifier ensemble composed of
weak classifiers trained with a boosting algorithm implemented in a
multiprocessor system on chip. The network is applied on the
classification on thyroid disease diagnosis. The objective is to show
that, even an FPGA with hardware restrictions, can be used to implement
a complex problem, when parallel processing is used. To improve the
system performance four soft processors were used with a shared memory.

#e1#

#s1#
#abstract1040.txt#


An FPGA-based custom core which computes the Gaussian calculation
portion of a Hidden Markov Model (HMM) based speech recognition system,
is presented. The work is part of the development of a custom embedded
system which will provide speaker independend, large vocabulary
continuos speech recognition and is currently presented as a
hardware/software codesign. By de-coupling the Gaussian calculation
from the backend search, calculation of Gaussian results is performed
with minimal communication between backend search software and an FPGA
based Gaussian core. Several implementations have been investigated in
order to minimize memory bandwidth and FPGA resource requirements and
are presented. The system has been implemented using an Alpha Data
XCR-5T1, reconfigurable computer housing a Virtex 5 SX95T FPGA and has
achieved better than real-time performance at 133MHz. The core has been
tested and is capable of calculating a full set of Gaussian results
from 3825 acoustic models in 5.3ms which coupled with a backend search
of 5000 words has provided over 80% accuracy.

#e1#

#s1#
#abstract1041.txt#


This paper describes a comparison of two FPGA Montgomery modular
multiplication architectures: a fully systolic array and a parallel
implementation. The modular multiplication is employed in modular
exponentiation processes, which is the most important operation of some
public-key cryptographic algorithms and the most popular of them is the
RSA encryption scheme. The proposed fully systolic array architecture
presents a high-radix implementation with carry propagation between the
Processing Elements. The parallel implementation is composed by
multipliers blocks in parallel with the Processing Elements and it
provides a pipelined operation mode. We compared the time x area
efficiency for both architectures as well as a RSA application. The
fully systolic array implementation can run the 1024 bit RSA decryption
process in just 3.23 ms and the parallel architecture executes the same
operation in 6 ms, which means a competitive state-of-art performance
for both architectures.

#e1#

#s1#
#abstract1042.txt#


Human-centric applications, like financial and commercial, depend on
decimal arithmetic since the results must match exactly those obtained
by human calculations. The IEEE-754 2008 standard for floating point
arithmetic has definitely recognized the importance of decimal for
computer arithmetic. A number of hardware approaches have already been
proposed for decimal arithmetic operations, including addition,
subtraction, multiplication and division. However, few efforts have
been done to develop decimal IP cores able to take advantage of the
binary multipliers available in most reconfigurable computing
architectures. In this paper, we analyze the tradeoffs involved in the
design of a parallel decimal multiplier, for decimal operands with 8
and 16 digits, using existent coarse-grained embedded binary arithmetic
blocks. The proposed circuits were implemented in a Xilinx Virtex 4
FPGA. The results indicate that the proposed parallel multipliers are
very competitive when compared to decimal multipliers implemented with
direct manipulation of BCD numbers.

#e1#

#s1#
#abstract1043.txt#


The lower hybrid (LH) power deposition and the current drive (CD)
efficiency were assessed by the application of modulated LH power.
Density and magnetic field scans were performed and the response of the
electron temperature provided by the available electron cyclotron
emission diagnostic was investigated by means of fast Fourier transform
analysis. An innovative technique based on a comparison between
modelled and experimental data was developed and used in the study. The
LH waves are absorbed by fast electrons with energies of a few times
the thermal one, causing a modification in the electron distribution
function (EDF) by creating a plateau in the parallel direction. The
phase of the temperature perturbations, phi , as well as the ratio
between the amplitudes of the third and the main harmonics, delta T/sub
e3// delta T/sub e1/, are found to be strongly affected by the plateau
of the EDF as the broader the plateau the larger | phi |, ( phi # 0),
and the smaller delta T/sub e3// delta T/sub e1/ are. Transport and
Fokker-Planck modelling was used to support this conclusion as well as
to interpret the experimental data and hence to assess the LHCD
efficiency and deposition profile. The results from the analysis are
consistent with broad off-axis LH power deposition profile. For
densities between 1 * 10/sup 19/ and 4 * 10/sup 19/ m/sup -3/, which is
the accessibility limit at the highest magnetic field discharges, a
gradual shift of the maximum of the power deposition to the periphery
and a degradation of the CD efficiency was observed.

#e1#

#s1#
#abstract1044.txt#


A high-speed one-dimensional detector for time-resolved small-angle
x-ray scattering has been designed and built for experiments at the
Advanced Photon Source of Argonne National Laboratory. This detector is
made from a 500- mu m thick by 150-mm diameter ultra-high-purity n-type
silicon wafer. The electrodes, which are a series of concentric rings
that are deposited in the wafer, integrate the scattered x-rays over
the azimuthal angle and, thereby, produce a one-dimensional detector.
This design yields 128 rings, which allows parallel processing of the
signal from each ring. The readout electronics consist of
transimpedance front-end amplifiers, one for each ring, followed by
active pulse-shaping filters. The amplifier signals are digitized using
12-bit analog-to-digital converters, one per ring, which operate at 20
MHz. The frame rate of the system is 271 kHz. Up to 2/sup 20/ - 1
scattering profiles may be stored on a random access memory chip and
transferred to a data file at a rate of 16 * 10/sup 3/ profiles/sec.
For X-ray energies between 3.5 and 13.2 keV the efficiency exceeds 80%.
The resolving time of the electronics is 300 ns, which is sufficient to
isolate electronically a single pulse of scattered x-rays when the
synchrotron is operated in a hybrid or asymmetric fill pattern.
Therefore, laser-pump/x-ray-probe experiments can be performed without
a mechanical shutter. Examples of time-resolved speckle and the
kinetics of the formation of sodium chloride particles are presented.
This detector is capable of acquiring small-angle x-ray scattering
profiles over multiple time scales, which are needed to characterize
many chemical, physical, and biological processes. In addition, this
detector may be tested and calibrated before experimental runs, without
access to an intense beam of x-rays, with alpha particles from a
radioactive source such as /sup 241/Am.

#e1#

#s1#
#abstract1045.txt#


This brief introduces a new method for sampling of transient analog
waveforms based on the parallel exponential filters. The signal is fed
to the parallel network consisting of resistor-capacitor (RC) circuits,
outputs of which are simultaneously sampled. We show that N previous
samples of the input signal can be reconstructed from single output
samples of N parallel RC circuits. The parallel sampling method
increases the sampling rate of the data acquisition system by a factor
of N. In particular, the method is useful in increasing the sampling
rate of the Flash-type analog-to-digital VLSI circuits. We present the
parallel RC network, develop the reconstruction algorithm, and briefly
describe a variety of applications such as measurement and
reconstruction of pulses produced by ultrawideband transmitters,
radiation detectors, and pulse lasers.

#e1#

#s1#
#abstract1046.txt#


This paper analyzes possible performance improvement of streaming
applications by the parallel computation platform of FPGAs. Software
developers still are not familiar with the hardware implementation
details of applications and will benefit from this analysis. First the
available logic and memory resources of modern FPGAs like Xilinx's
Virtex-5 and Virtex-6 devices are explored to determine how many
parallel processors can be configured with the available logic
resources and how many parallel processors' speed can be sustained by
simultaneous data access from the on-chip memory. A portion of the
on-chip memory is first set aside for pre-fetching of reconfiguration
bits for dynamic reconfiguration of FPGA. The rest of the on- chip
memory is used for data memory requirements of the concurrent
processors. Based on input data rate and required output throughput of
data, a number of on-chip cache memory locations and their sizes are
determined. Finally a quantitative analysis of possible performance
gain of streaming applications by FPGA implementation is compared to
that of sequential implementations by a 2.5 GHz dual-core Intel
microprocessor.

#e1#

#s1#
#abstract1047.txt#


Petersen graph has good performance in parallel and distributed
computation because of its properties such as short diameter and
regularity. Based on the simple scalable properties of ring and the
short network diameter of Petersen graph, a new interconnection network
RP/sub n/(k) is proposed and the properties of the RP/sub n/(k) are
analyzed. It is proved that RP/sub n/(k) not only has good regularity
and extensibility, but also has shorter diameter and smaller
construction costs than the RP(k) network. Additionally, The conditions
satisfying that the network diameter of RP/sub n/(k) are better than
those of 2D Torus are presented.

#e1#

#s1#
#abstract1048.txt#


Computing systems typically suffer from delay in data processing. This
delay is caused by computational power, architecture of the processor
unit, synchronization signals, and so on. To enhance the performance of
these systems by increasing the processing power, a new architecture
and clocking technique is carried out in this paper. This new
architecture design called Embedded Parallel Systolic Filters (EPSF)
that can process data gathered from sensors and landmarks are proposed
in our study using a high-density reconfigurable device (FPGA chip).
The results show that EPSF architecture and bit-flag with a flicker
clock perform significantly better in multiple input sensors signals
under both continuous and interrupted conditions. Unlike the usual
processing units in previous tracking and navigation systems used in
robots, this system allows autonomous control of the robot through a
multiple technique of filtering and processing. Furthermore, it
provides fast performance and a minimal size for the entire system that
minimizing the delay about 70%.

#e1#

#s1#
#abstract1049.txt#


Multi-cores are the contemporary solution to satisfy high performance
and low energy demands in general and embedded computing domains.
However, currently available multi-cores are not feasible to be used in
safety-critical environments with hard real-time constraints. Hard
real-time tasks running on different cores must be executed in
isolation or their interferences must be time-bounded. Thus, new
requirements also arise for a real-time operating system (RTOS), in
particular if the parallel execution of hard real-time applications
should be supported. In this paper we focus on the MERASA system
software as an RTOS developed on top of the MERASA multi-core
processor. The MERASA system software fulfils the requirements for
time-bounded execution of parallel hard real-time tasks. In particular
we focus on thread control with synchronisation mechanisms, memory
management and resource management requirements. Our evaluations show
that all system software functions are time-bounded by a worst-case
execution time (WCET) analysis.

#e1#

#s1#
#abstract1050.txt#


In this paper, we report a fast implementation of Wyner-Ziv video
decoder using general-purpose computing on graphics processing units
(GPGPU). Despite of its many advantages, Wyner-Ziv video coding has a
problem of huge decoding complexity. Since Slepian-Wolf decoding with
rate adaptive LDPC accumulate code takes up more than 90% of entire
Wyner-Ziv video decoding complexity, in this paper, we focus on fast
implementation of the Slepian-Wolf decoder using the CUDA (Compute
Unified Device Architecture) which is a GPGPU architecture developed by
NVIDIA. Our implementation is shown to be 4~5 times (QCIF size) or
15~20 times (CIF size) faster compared to conventional Slepian-Wolf
decoding.

#e1#

#s1#
#abstract1051.txt#


A new kind of Intrusion Detection System (IDS) based on Principal
Component Analysis (PCA) and Grey Neural Networks (GNN) is presented to
improve the performance of BP neural networks in the field of intrusion
detection. First, the pre-processed data set is normalized and the
features of them are extracted by PCA. Next, five layers of the grey
neural networks is designed based on BP neural networks and Grey
theory, then the IDS composed of sniffer module, data processing
module, grey neural network module and intrusion detection module is
presented. Finally, the presented system was tested on the data set of
DARPA 1999. The results demonstrate that the feature extraction reduced
the dimensionality of feature space greatly without degrading the
systems' performance, and GNN not only promote the parallel computing
power of the system but also improve the utilization of available
information.

#e1#

#s1#
#abstract1052.txt#


As the wide application of multi-core processor architecture in the
domain of high performance computing, fault tolerance for shared memory
parallel programs becomes a hot spot of research. For years,
checkpointing has been the dominant fault tolerance technology in this
field, and recently, many research works have been engaged with it.
However, to those programs which deal with large amount of data,
checkpointing may induce massive I/O transfer, which will adversely
affect scalability. To deal with such a problem, this paper proposes a
fault tolerance approach, making use of redundancy, for shared memory
parallel programs. Our scheme avoids saving and restoring computational
state during the program's execution, hence does not involve I/O
operations, so presents explicit advantage over checkpointing in
scalability. In this paper, we introduce our approach and the related
compiler tool in detail, and give the experimental evaluation result.

#e1#

#s1#
#abstract1053.txt#


Distributed network system based on Controller Area Network (CAN)
protocol has been widely applied into the vehicle industries. CAN
communication protocol for the parallel-serial hybrid electrical
vehicles (PSHEV) has been designed in the paper. Real-time scheduling
algorithms based on optimization periodic algorithm (OPA) with
non-preemptive and preemptive scheduling are also proposed with an
optimization goal function to decrease the delay time for sporadic
tasks and increase the CPU utilization for a real-time controller. Then
a kind of method of solving the OPA have been designed namely exhaust
algorithm (EA). With a view to avoid real vehicles experiment failures
in advance, it is necessary for us to make an experiment set-up to
verify the validity of the CAN protocol designed and OPA for the PSHEV.
So the related hardware circuits for testing CAN system have been made.
Finally, communication experiments are made based on CAN protocol
proposed above. The results show the CAN protocol and OPA scheduling
algorithm for the PSHEV are valid, all messages can be transmitted and
received efficiently and completely avoiding the occurrence of loss of
messages in the process of communications in all nodes of the system
and the optimized periodic tasks as a result of OPA scheduling
algorithm can decrease the delay time for sporadic tasks and increase
the CPU utilization for a real-time controller.

#e1#

#s1#
#abstract1054.txt#


We have constructed a virtual machine PC cluster that uses a virtual
machine as worker nodes, and proposed a method that acquires
insufficient resources dynamically from cloud computing systems while
basic computation is performed on its own local clusters. Virtual
machine environments enable us to manage computer resources flexibly
and make use of migration. For data access, we have used iSCSI protocol
which supports to access data through IP networks, and migrated virtual
machine to cloud where data is stored over a high latency network. We
have confirmed that the execution time becomes shorter with the
migration of virtual machines even the cost of migration is taken into
account, compared with the case of accessing data over a network when
an I/O-intensive application is executed.

#e1#

#s1#
#abstract1055.txt#


Recent technological advances in Micro Electro Mechanical Systems
(MEMS) have enabled the design of lowcost, lightweight sensor nodes
capable of sensing, processing and communicating different types of
data. These tiny sensor nodes leverage the ideas found in Wireless
Sensor Networks (W

#e1#

#s1#
#abstract1056.txt#


This article highlights the Nanchang University, Chang-Da on the 1st
series and parallel using a group of 3.3V Li-ion battery boxes made of
288V, to a three-phase brushless DC motor to drive, and by
TMS320LF2407A DSP to control the PWM output, and then adopted by the
IGBT device composed of three-phase full-bridge circuit to drive the
brushless DC motor speed, thus ensuring the motor speed control
requirements, and also designed a by TL431, PC817, HCPL316J, UC3844,
LT1117 chips such as the composition of the DC / DC conversion.Through
this circuit, you can accurately get the output of +/-12V, +/-15V,
+/-9V and 3.3V voltage, so that the these voltages can give electric
cars the various electronic devices, as well as DSP power to maintain
the normal operation of electric vehicles by Chang Da on the 1st
successful development of electric vehicles, indicating the reliability
of the design.

#e1#

#s1#
#abstract1057.txt#


Most of embedded multiprocessor platforms are ideal for running diverse
operating systems and implementing different applications.
Inter-processor communication interface makes it possible for an
embedded multi-processor system to easily support multiple subsystems
parallel processing. This paper propose a kind of OS-level
communication interface implementing method based-on HPI in the
Complementary Multi-processor System and present the primary principles
and corresponding DSP side code structure for HPI. Take a single-board
embedded multimedia system which was integrated three embedded
microprocessors (PXA255, TMS320DM642 and SM501) as hardware platform,
detailed description of how implement tasks communication and data
transform by HPI between different embedded OS(ARM-Linux and mu
C/OS-II) running on ARM and DSP separately.

#e1#

#s1#
#abstract1058.txt#


The Variable Preconditioned GVR (VPGCR) with mixed precision on
Graphics Processing Unit (GPU) using Compute Unified Device
Architecture (CUDA) is numerically investigated. The convergence
theorem of VPGCR is guaranteed that the residual equation for the
preconditioned procedure can be solved in the range of single precision
operation. The results of computations show that VPGCR with mixed
precision operation on GPU demonstrated significant achievement than
that of CPU. Especially, VPGCR on GPU with mixed precision operation is
22.53 times faster than that of Central Processing Unit (CPU).

#e1#

#s1#
#abstract1059.txt#


The classical solution of electromagnetic problems using the finite
element (FE) method needs to assemble, store and solve an Ax = b matrix
system. A new technique for solving FE cases, considered much simpler
than traditional methods, shows that the assembling of the matrix A is
unnecessary. The difference between these two techniques is the
computation and processing time. The new one requires more iterations
to converge, observing, nevertheless, that the results are reliable.
One possible way to improve its performance is the application of
parallelization techniques.

#e1#

#s1#
#abstract1060.txt#


Parallel computation is application-oriented, particularly for the GPU
(Graphics Processing Unit) with the inherent parallelism. This paper
shows the architecture of a GPU cluster based on MPI (Message Passing
Interface) and CUDA (Compute Unified Device Architecture). Results show
that the acceleration ratio is obviously improved but the acceleration
effect seems decelerated in large-scale GPU cluster. The parallel
algorithm is mainly focused on task partitioning sparse matrix-vector
multiplications (SpVM) in GPUs.

#e1#

#s1#
#abstract1061.txt#


Today's low cost hardware developments allow for a parallelization of
intensive computation processes over several selectable CPU cores by
using Intel's OpenMP library. But if one applies this feature to a
numerical simulation based on a boundary element method (BEM), there is
still a huge bottle neck - the shared memory of such a system. So, if a
low cost hardware should be effectively used for large BEM problems, a
memory compression algorithm that is easily to be scheduled in parallel
is of first choice. Within our new idea and development of the
Hierarchical Block Wavelet Compression (HWC) which is based on IEEE's
JPEG2000 standard for image compression, this bottle neck will be
tackled in pure mathematically manners. Furthermore, its
parallelization will be discussed and an optimal compression rate for a
3-D electrostatic BEM problem by a rather simple optimization algorithm
will be presented.

#e1#

#s1#
#abstract1062.txt#


We report a novel application of a graphics processing unit (GPU) for
the purpose of accelerating the search pipelines for gravitational
waves from coalescing binaries of compact objects. A speed-up of
16-fold in total has been achieved with an NVIDIA GeForce 8800 Ultra
GPU card compared with one core of a 2.5 GHz Intel Q9300 central
processing unit (CPU). We show that substantial improvements are
possible and discuss the reduction in CPU count required for the
detection of inspiral sources afforded by the use of GPUs.

#e1#

#s1#
#abstract1063.txt#


This paper describes several practical steps for accurate statistical
modeling of a known acoustical noise environment to attain good
performance of a small vocabulary speech recognizer for isolated words
based on whole-word hidden Markov models. Hierarchical segmentation
based on Bayes information criterion and k-means clustering followed by
split-merge Gaussian mixture model training were utilized for noise
model estimation. Parallel model combination technique produces final
noise-corrupted speech models for a small group of speakers.
Experiments were carried out on a real operating room ambient noise
recorded during a neurosurgery at the University Hospital in Marburg.

#e1#

#s1#
#abstract1064.txt#


The long computation time caused by sequential processing of huge size
image in three-dimensional image direct writing prevents its
industrialization. To reduce computation time, OpenMP technique is
applied to process the images in a parallel way. Performance improves
by both parallel image reading and parallel image processing.
Experiment results show that the method possesses high effectiveness
and reliability.

#e1#

#s1#
#abstract1066.txt#


Autonomous signal detection of the North Atlantic right whale (NRW),
Eubalaena glacialis, is becoming an important factor in monitoring and
conservation for this highly endangered species. Both online and
offline systems exist to help study and protect animals within this
population. In both cases auto-detection of species-specific calls
plays a vital role in localizing individual animal by searching
time-frequency passive acoustic data. This research presents an
experimental system, referred to as the NRW-CRITIC, for automatic
detection of the NRW contact call. In general, the CRITIC uses a
combinatorial classifier approach to integrate a series of existing
machine learning algorithms; each designed specifically for NRW contact
call identification. The proposed configuration consists of several
recognition methods running in parallel; these include linear
discriminant analysis, artificial neural network (NET) and
classification regression tree (CART). This paper presents the details
for the NRW-CRITIC and discusses the approach used to combine multiple
independent decisions into a single result. A side-by-side performance
comparison, between the CRITIC and a well-known method, the feature
vector testing (FVT), is summarized. Performance metrics are evaluated
based on a large database of acoustic recordings consisting of over
58,000 NRW contact calls from various locations, including two critical
habitats, Great South Channel and Cape Cod Bay. Results indicate the
FVT algorithm yields a 74.7% detection probability with an error rate
of 4.35%. In comparison the CRITIC, operating at similar information
level yields a 78.02% detection probability with a 3.25% error rate,
exceeding the performance of the FVT. Performance was also measured
using data from a multi-channel acoustic array located in Massachusetts
Bay. A side-by-side comparison of array presence is discussed for two
separate days. Results show that with the FVT and CRITIC operating at
0% error for array presence, the FVT method had 18,769 and 24,469 false
positives for the Massachusetts Bay datasets respectively. With the
same 0% error condition the CRITIC provided successful detection with
significantly lower number of false positive rates: 1,072 and 2,324
calls, respectively. Future extensions of this experimental work are
also discussed.

#e1#

#s1#
#abstract1067.txt#


We have previously demonstrated a fabrication technique for the
creation of silicon wire grid polarizers (WGPs) for far (deep)
ultraviolet applications utilizing a shear-aligned cylinder-forming
polystyrene-b-poly(n-hexyl methacrylate) diblock copolymer as a mask
for reactive ion etching of an amorphous silicon substrate. In our
current work, a numerical model is refined and applied to our
experimental systems to describe the impact of wire height and
periodicity, and tradeoffs between the two, on polarization efficiency.
We focus our attention at a wavelength of 193 nm, the emission
wavelength of the ArF excimer laser currently in use in advanced
photolithographic processes. Through application of the model's
predictions we have achieved marked improvement in the polarization
efficiency of our WGPs by increasing the block copolymer molecular
weight, thereby increasing the thickness of the Si wires, which
compensates for a simultaneous increase in wire periodicity; the
resulting arrays of parallel Si nanowires exhibit polarization
efficiencies approaching 64% at 193 nm, a 68% relative increase over
our previous Si WGPs.

#e1#

#s1#
#abstract1068.txt#


Amorphous hydrogenated silicon nitride (SiNH) materials prepared by
plasma-enhanced chemical vapor deposition (PECVD) are of high interest
because of their suitability for diverse applications including optical
coatings, gas/vapor permeation barriers, corrosion resistant, and
protective coatings and numerous others. In addition, they are very
suitable for structurally graded systems such as those with a graded
refractive index. In parallel, modeling the PECVD process of SiN(H) of
an a priori given SiN(H) ratio by atomistic calculations represents a
challenge due to: (1) different (and far from constant) sticking
coefficients of individual elements, and (2) expected formation of
N/sub 2/ (and H/sub 2/) gas molecules. In the present work, we report
molecular-dynamics simulations of particle-by-particle deposition
process of SiNH films from SiH/sub x/ and N radicals. We observe
formation of a mixed zone (damaged layer) in the initial stages of film
growth, and (under certain conditions) formation of nanopores in the
film bulk. We investigate the effect of various PECVD process
parameters (ion energy, composition of the SiH/sub x/+N particle flux,
ion fraction in the particle flux, composition of the SiH/sub x/
radicals, angle of incidence of the particle flux) on both (1)
deposition characteristics, such as sticking coefficients, and (2)
material characteristics, such as dimension of the nanopores formed.
The results provide detailed insight into the complex relationships
between these process parameters and the characteristics of the
deposited SiNH materials and exhibit an excellent agreement with the
experimentally observed results.

#e1#

#s1#
#abstract1069.txt#


Staged design has been introduced as a programming paradigm to
implement high performance Internet services that avoids the pitfalls
related to conventional concurrency models. However, this design
presents challenges concerning resource allocation to the individual
stages, which have different demands that change during execution. On
the other hand, processing resources have been shown to form the
bottleneck in a variety of Internet-based applications. For this
reason, parallel processing hardware techniques have been employed in
order to cope with the massive concurrency and the increasing demands
for performance aspects in these applications. Recently, the rise of
multi-core technology introduces a hierarchic parallelism in modern
server machines that has to be considered when allocating processing
units in order to improve the utilization of these resources. This
paper, introduces an adaptive policy to allocate processing units in
Internet services that are based on the staged architecture. The
proposed approach takes the hierarchic parallelism in account and
adapts the resources assigned to the individual stages dynamically
based on the observed demand of each stage using a feedback loop.
Simulation results demonstrate that our approach achieves a competitive
system throughput, avoids overhead that is associated with parallel
processing, and successfully adapts resource allocation to dynamic
changes in workload characteristics.

#e1#

#s1#
#abstract1070.txt#


Recently, hardware and software engineers have been showing
considerable attention to high-level parallelization and hardware
synthesis methodologies. State-of-the-art approaches have benefited
from the emergence of modern high-density Field Programmable Gate
Arrays. In this paper, we explore the effectiveness of a formal
methodology in the design of pipelined versions of a matrix
multiplication algorithm. The suggested methodology adopts a functional
programming notation for specifying algorithms and for reasoning about
them. The parallel behavior of the specification is then derived and
mapped onto hardware. Several pipelined implementations are developed
with different performance characteristics. The refined designs are
tested under Agility's RC-1000 reconfigurable computer with its 2
million gates Virtex-E FPGA. Performance analysis and evaluation of the
proposed implementations are presented in comparison with an Intel Core
2 DUO processor.

#e1#

#s1#
#abstract1071.txt#


We use a 3-D parallel processing software based on the FDTD method to
efficiently simulate the EM problem with an ill-conditioned
computational domain. Advanced PML boundary condition allows us to use
several cells of white space between the object and domain boundary and
can generate the accurate results. A network card and a PEC sheet with
a special slot are used to verified the proposed technique for the ECC
(envelope Correlation Coefficient) and the transmitted power
calculation. The simulation results match well with the other methods
and measurement data.

#e1#

#s1#
#abstract1072.txt#


A bidimensional pixel CMOS detector array for parallel in-pixel
demodulation and time-resolved correlation computation is presented.
Optical signals can be processed at a rate higher than 10 000 samples
per second with demodulation frequencies in the megahertz range.

#e1#

#s1#
#abstract1073.txt#


Silicon nanograss and nanostructures are realized using a modified deep
reactive ion etching technique on both plane and vertical surfaces of a
silicon substrate. The etching process is based on a sequential
passivation and etching cycle, and it can be adjusted to achieve
grassless high aspect ratio features as well as grass-full surfaces.
The incorporation of nanostructures onto vertically placed parallel
fingers of an interdigital capacitive accelerometer increases the total
capacitance from 0.45 to 30 pF. Vertical structures with features below
100 nm have been realized.

#e1#

#s1#
#abstract1074.txt#


Algorithms for real-time parallel processing of an audio signal in
large-scale digital audio distribution networks, implemented on
personal computer platform, are compared in this paper from the
performance point of view. In such systems, the summing and
multiplication of audio signal take up a significant portion of the
processing power of the system. Therefore, there is a tendency to
decrease its computing demands and thus reserve the computing power of
the processor for other signal processing modules in the audio network.
Various approaches can therefore be used to distribute computing among
several threads. Seven approaches are analyzed in the paper and fastest
one is found with regard to the size of audio sample buffers.

#e1#

#s1#
#abstract1075.txt#


The paper investigates an impact of direct and combining collective
communications models that may be critical for performance of parallel
applications. Analysis provided for any given start-up time and message
transfer time reveals the fastest collective communication mode in
relation to the number of processing elements in 2D meshes and fat tree
networks on a chip.

#e1#

#s1#
#abstract1076.txt#


This study describes the performance results on testing MatLab
applications using the parallel computing and the distributed computing
toolboxes under different platforms with different hardware and
operating systems. Each trial was executed keeping the hardware fixed
and changing the operating system to obtain unbiased results. To
standardize the benchmarking test, Fast Fourier Transform (FFT),
discrete cosine transform (DCT), edge detection and matrix
multiplication algorithms were executed. The results show that the
leveraging of multicore platforms can speed up considerably the
processing of images through the use of parallel computing tools in
MatLab. Two different system hardware platforms (systems 1 and 2) were
used in a series of experiments. Four rounds of experiments were
performed benchmarking the FFT algorithm using the parallel tool box,
by changing system platform, number of workers, image size and number
of images. The results of the ANOVA test suggest that although there is
no statistical significance on the factor represented by the operating
system (OS) on system 1, the OS plays a significant roll on system 2.
Moreover, on both systems there is statistical significance on the
factors represented by the number of workers utilized and the number of
images processed, yielding more than a 500% performance increase by
using 8 MatLab workers on a dual quad-core machine.

#e1#

#s1#
#abstract1077.txt#


Parallel and distributed systems have intended to exploit local and
remote multi-cores and multiprocessors for high performance. However,
without popular single-image distributed operating systems, few systems
can utilize all resources automatically and effectively. This paper
proposed an MPI-like middleware, Pitcher, to distribute multithreaded
parallel workloads across networked computers without user involvement.
Fine-grained computational units, threads, are regrouped into bundles
based on data locality and scheduling requests. Such thread bundles are
treated as load distribution units. Unlike MPI, data and
synchronization variables in Pitcher will be automatically partitioned
and distributed. A preprocessor transforms the source code so that
programmers stick to shared virtual address space programming paradigm
whereas the runtime support module dispatches thread bundles and data
based on locality. The experimental results demonstrate the
effectiveness of Pitcher.

#e1#

#s1#
#abstract1078.txt#


MapReduce is a programming framework introduced by Google for
large-scale data processing. It is usually used in a scan-centric
fashion where all the data are split into blocks and Maps are generated
for each block to scan and process the data in the block, then Reduces
merge outputs from all the Maps. When a query intends to process only a
subset of the data selected by a predicate, this brute-force method may
cause extra I/O overhead spent on irrelevant data, and the overhead for
initiating so many Maps may be non-trivial given that the actually
interesting data for the query is comparatively small in volume. We
propose an approach to integrate the index into the MapReduce execution
in which only an appropriate number of Maps are generated, each of
which accesses the data using an index. This approach incurs random I/O
and remote access to data, so the overall performance depends on both
system parameters and the query characteristics. We build a cost model
for both this index access execution and the traditional full scan
execution. This cost model can be used to choose between the two
execution modes before executing a query. Experiments show that the
index access execution can greatly outperform full scan execution when
the selectivity of the predicate is low, and the cost model predicts
the actual execution cost very well so can be used to determine the
execution plan for a query.

#e1#

#s1#
#abstract1079.txt#


The requirements of OLAP applications increase rapidly by dramatically
increased data volume, users, query volume and query complexity. The
requirement for shortening update period in data warehouse is another
crucial factor for a scalable OLAP application. In this paper, we
propose a scalable OLAP prototype to support the query processing with
increasing data volume by distributing the whole fact tuples to
multiple servers to construct a set of sibling cubes which can be
merged together to obtain the whole cube. We employ a light weight
distribution policy with fully duplicated dimension tables in each
sibling server on the observation of very low proportion of space cost
for dimension tables. OLAP query with distributed aggregate functions
can be transformed into queries to be performed parallel in sibling
servers. For non-distributed computing aggregate functions, such as
median, the optimized median aggregate computing algorithm is proposed
to reduce transmission volume between servers while computing the
global median values. We also present a three-level framework in data
warehouse to meet the requirement of shorter update period in
"operational business intelligence". An asynchronous tunnel model is
proposed to reduce update latency by pre-fetching updated tuples to
OLAP processing server. Finally, we set up prototype system ParaCube to
evaluate performance in

#e1#

#s1#
#abstract1080.txt#


A model for image and data pre-processing and communication between a
dedicated PC and a PLC with Neural Network (NN) application is proposed
in this paper. The proposed model defines guidelines for creating a
multithreaded application for receiving real-time data from several
digital cameras, parallel image pre-processing based on predefined user
algorithms, calculation of input data vector for NN and sending the
input vector to the PLC NN application. The model was developed and
verified in the laboratory "Intelligent Manufacturing Systems" at the
Technical University of Sofia.

#e1#

#s1#
#abstract1081.txt#


Modern FPGA chips, with their larger memory capacity and
reconfigurability potential, are opening new frontiers in rapid
prototyping of embedded systems. With the advent of high density FPGAs
it is now possible to implement a high performance VLIW processor core
in an FPGA. Architecture based on Very Long Instruction Word (VLIW)
processors are an optimal choice in the attempt to obtain high
performance level in embedded system. In VLIW architecture, the
effectiveness of these processors depends on the ability of compilers
to provide sufficient instruction level parallelism(ILP) in program
code. Using advanced compiler technology could take these functions,
This paper describes research result about enabling the DSP TMS320
C6201 model that be described with machine description language (MDES)
in compiler technology for image processing applications by exploiting
FPGA technology and assembly code that be more known as Lcode would be
generated by the compiler depends on MDES given when running the
compiler. We present a DSP C6201 VHDL from MDES definition with VLIW
architecture model using compiler technology. We call this new
development as Modified Minimum Mandatory Modules (M4) approach that be
derived from M3 methodology. Our goals are to keep the flexibility of
DSP in order to shorten the development cycle. Our results demonstrate
that an algorithm can easily, in an optimal manner, specified and then
converted to VHDL language and implemented on an FPGA device with
system level software. This makes our approach suitable for developing
co-design environments. Our approach applies some criteria for
co-design tools : flexibility modularity, performance, and reusability.

#e1#

#s1#
#abstract1082.txt#


Optical flow is a motion field estimation method that has a wide range
of applications. In this paper, we present a fully pipelined hardware
architecture for high-speed optical flow estimation based on a
full-search block matching algorithm. A census transform is applied to
the corresponding pixels in the current and previous frame. The
similarity between two census vectors within the search area is then
computed by measuring the hamming distance. Macro blocks are generated
based on the measured hamming distance values and the best match is
determined by locating the block that has the smallest sum. The
synthesis tool reported that the proposed system is capable of
processing 400 standard VGA frames per second.

#e1#

#s1#
#abstract1083.txt#


Multi-field Packet classification is the main function in
high-performance routers. The current router design goal of achieving a
throughput higher than 40 Gbps and supporting large rule sets
simultaneously is difficult to be fulfilled by software approaches. In
this paper, a set pruning trie based pipelined architecture called Set
Pruning Multi-Bit Trie (SPMT) is proposed for multi-field packet
classification. However, the problem of rule duplications in SPMT that
may cause a memory blowup must be solved in order to implement SPMT
with large rule sets in FPGA devices consisting of limited on-chip
memory. We will propose two rule grouping schemes to reduce rule
duplications in SPMT. The first scheme called Partition by Wildcards
(PW) divides the rules into subgroups based on the positions of their
wildcard fields. The second scheme called Partition by Length (PL)
rules partitions the rules into subgroups according to their prefix
lengths. Based on our performance experiments on Xilinx Virtex-5 FPGA
device, the proposed pipeline architecture can achieve a throughput of
over 100 Gbps with dual port memory. Also, the rule sets of up to 10k
rules can be fit into the on-chip memory of Xilinx Virtex-5 FPGA device.

#e1#

#s1#
#abstract1084.txt#


This paper proposes a design of a database system which accelerates the
execution of database transactions by offloading database operators
(e.g. joins, scans and sorting) in hardware algorithms executed on
runtime reconfigurable computing platforms. Furthermore, a hybrid
database system is described which exploits the strengths of the new
reconfigurable hardware-based database system in combination with
pre-existing technologies for main memory or disc resident database
systems. Due to the parallel and pipelined design style of the hardware
algorithms, the new database system offers a potential speedup over the
traditional sequential execution on instruction stream processors.
Moreover, data access times can be reduced by placing circuits on the
reconfigurable fabric close to embedded memory which further allows for
customizing high bandwidth memory interfaces. As a consequence of the
accelerated execution speed, the new system has the potential to
reliably meet constraints in real-time scenarios, to reduce the chance
of lock contention and cache flushes, and to decrease the cost for
concurrency control.

#e1#

#s1#
#abstract1085.txt#


We present a GPU-accelerated algorithm for plant growth modeling and
visualization. In this algorithm, plant topological structures and
geometrical structures are represented separately. Plant topological
structures are generated by dual-scale automaton, and geometrical
structures are dynamically constructed in parallel by the geometry
shader of Graphics Processing Unit (GPU). This scheme greatly reduces
the amount of transferred data between main memory and GPU memory, and
speeds up plant growth modeling and rendering without compromising
image quality. An application for cotton growth modeling and
visualization is given in detail to confirm the advantages of the
proposed algorithm.

#e1#

#s1#
#abstract1086.txt#


The imaging of bistatic SAR with asynchronous transceiver positions is
studied in this paper. First of all, the geometry of the asynchronous
transceiver model is constructed, and the approximation model of the
echo from the bistatic SAR is obtained based on the analysis of the
instantaneous transceiver distance. Asynchronous transceiver positions
will lead to non-uniform SAR data. This paper first analyzes the method
for transforming the asynchronous bistatic model into a monostatic
variable motion model. Then with the non-uniform FFT, it resolves the
non-uniform sampling problems induced by nonconstant velocities. The
simulation experiments show that the bistatic SAR images have good
quality, which verifies the validity of the method presented in this
paper.

#e1#

#s1#
#abstract1087.txt#


Open Computing Language (OpenCL) is a fundamental technology for
cross-platform parallel programming. The emerging of OpenCL provides
portable and efficient access to the power of modern processors. This
revolutionary new technology is applied to accelerate the
reconstruction of cone beam computed tomography (CBCT) on Graphics
Processing Unit (GPU) in this paper. An OpenCL-based implementation of
the Feldkamp-Davis-Kress (FDK) algorithm is presented. The required
transformations to parallelize the algorithm for the OpenCL
architecture are also explained. Comparing to the conventional
CPU-based implementation, the proposed method reaches an over 57 times
speedup. Experimental results show a great performance boost, which can
pave the way for widespread application and new conceptual innovation
of CBCT. Besides, the feasibility and potential of OpenCL-based
implementation are also indicated.

#e1#

#s1#
#abstract1088.txt#


This paper describes a new implemented method for the MPEG audio layer
III (MP3) decoder. The proposed architecture is based on a graphic
process unit (GPU) using CUDA environment, where it can effectively
take advantage of modern GPU's parallel computing power. The
implemented system with this architecture employs a multi-thread model
and memory optimization to process MP3 decoding in parallel, so it is
significant to minimize the computational overhead. Experimental
results on a GTX260+ graphics card showed that the proposed
architecture is over five times faster than traditional MP3 library
based on CPU.

#e1#

#s1#
#abstract1089.txt#


Block Truncation Coding (BTC) is an efficient compression technique for
its inherent simple coding strategy. However, the annoying blocking
effect and false contour accompanied in high coding gain configurations
make the applications relatively limited compares to some up-to-date
compression schemes. For this, Error-Diffused Block Truncation Coding
(EDBTC) is proposed to solve these problems and obtain satisfactory
results. Unfortunately, the EDBTC sacrifices the parallel advantage of
traditional BTC. Moreover, the number of diffused directions of EDBTC
can be reduced to obtain higher efficiency. For these, the Interlaced
Error-Diffused Block Truncation Coding (IEDBTC) is proposed in this
work to claim back the parallel advantage. In addition, the diffused
elements are also reduced from four to two with the proposed
optimization procedure while preserving the image quality.

#e1#

#s1#
#abstract1090.txt#


In the past, efforts to speed up motion estimation for video encoding
were directed at finding better predictive search algorithms. Now, they
are directed toward the shrewd exploitation of the machine's advanced
architectural features such as multimedia extensions, especially for
the computation of the error metric which is known to be expensive. In
this paper, we extend previous work by further exploring efficient
implementation of approximate fast metrics for motion estimation. We
show that the proposed metrics can be implemented using SIMD
instructions to yield impressive speed-ups, up to 12:1 relative to
non-vectorized but otherwise optimized C code, while sacrificing less
than 0.1 dB on image quality.

#e1#

#s1#
#abstract1091.txt#


Embedded microprocessors require efficient supply management systems to
optimize its power consumption and to enhance their calculation
potentials. Typically, the modules performing this function are known
as Voltage Regulator Modules (VRMs). It is widely adapted that
current-programmed regulation techniques own leveraging skills in the
control of this kind of power converters. However, these strategies
require a fine inductor-current sensing to achieve accurate results.
One critical issue in the inductor-current sensing is the effect of
parasitic inductances in the measurement loop. This undesirable effect
produces a considerable mismatch between the real inductor-current
waveform and the equivalent voltage image captured thanks to the shunt
resistance. Further, this unwanted deviation augments as long as the
current value is increased. As a result, this problem makes loosely the
data obtained. However, today's commercial digital controllers, like
FPGAs/sup 1/, can be used to reduce overwhelmingly the aforementioned
drawback. The presented work exploits some intrinsic advantages of
FPGAs such as its great processing speed and its parallel working mode
to overcome this drawback. Therefore, a new digital auto-tuning system
is proposed in which this undesirable effect is treated and
compensated. The obtained result is a digital signal which avoids the
parasitic effect of the inductance in the measurement loop. In the last
part of our work, some experimental results, using a FPGA, validate the
advantages of the proposed method.

#e1#

#s1#
#abstract1092.txt#


A transaction manager integrated directly into a field device prevents
inconsistencies to the data of the field device caused by parallel
accesses by more than one user. A transaction manager controls all
internal and external accesses to the field device data. Its functions
must be available not only within the field device but especially to
external users accessing the field device by standard industrial
communication systems. Different communication paradigms applied by
these communication systems require elaborated mapping approaches which
must be standardized in future.

#e1#

#s1#
#abstract1093.txt#


In this paper, we introduce an intrinsically parallel framework
striving for increased flexibility in development of robotic, computer
vision, and machine intelligence applications. The primary goal is to
provide a sound and easy-to-use, but yet efficient base architecture
for complex sensor-based robotic systems with focus on industrial
scenarios. The framework combines promising ideas of recent
neuroscientific research with a blackboard information storage
mechanism and an implementation of the multi-agent paradigm.
Additionally, a generic set of tools for realtime data acquisition and
robot control, integration of external software components, and user
interaction is provided. The paper is completed with a tutorial section
showing how the building blocks afore described can be composed to
applications of increasing complexity.

#e1#

#s1#
#abstract1094.txt#


To make the real-time focusing possible in the exposure process,
diffractive microlens arrays with continuous relief are designed and
fabricated using harmonic diffraction theory for parallel laser direct
writing to integrate the exposing and autofocusing functions in one
array by taking both the writing resolution and diffraction efficiency
into consideration. A theoretical model is established using
Rayleigh-Sommerfeld diffraction theory to accurately characterize the
focusing characteristics of each harmonic diffractive microlens in the
array so that the fidelity of pattern can be improved through exposure
dose modulation. The measurements made indicate that the experimental
results coincide well with the theoretical results when the writing
laser with a wavelength of 441.6 nm and the autofocusing laser with a
wavelength of 670 nm are normally incident on an array with an F-number
of F/4 fabricated on fused silica, and the array developed can be used
to synchronously focus the writing laser and the autofocusing laser
into the same spot of the array. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1113.txt#


Multicore architectures have established themselves as the new
generation of computer architectures. As part of the one core to many
cores evolution, memory access mechanisms have advanced rapidly.
Several new memory access mechanisms have been implemented in many
modern commodity multicore architectures. By specifying how processing
cores access shared memory, memory access mechanisms directly influence
the synchronization capabilities of multicore architectures. Therefore,
it is crucial to investigate the synchronization power of these new
memory access mechanisms. This paper investigates the synchronization
power of coalesced memory accesses, a family of memory access
mechanisms introduced in recent large multicore architectures such as
the Compute Unified Device Architecture (CUDA). We first define three
memory access models to capture the fundamental features of the new
memory access mechanisms. Subsequently, we prove the exact
synchronization power of these models in terms of their consensus
numbers. These tight results show that the coalesced memory access
mechanisms can facilitate strong synchronization between the threads of
multicore architectures, without the need of synchronization primitives
other than reads and writes. In the case of the contemporary CUDA
processors, our results imply that the coalesced memory access
mechanisms have consensus numbers up to 64.

#e1#

#s1#
#abstract1114.txt#


Fe-Al/sub 2/O/sub 3/ films have been prepared by on-off deposition of
Fe with continuous Al/sub 2/O/sub 3/ deposition, and also subsequently
irradiated by 400 keV Ar ions for controlling of the structural and
magnetic properties of the films. CEM spectra of the as-deposited films
show the decrease in the average hyperfine fields (H/sub hfs/) with
decreasing Fe/ Fe-Al/sub 2/O/sub 3/ volume ratio with keeping the H/sub
hfs/ orientation parallel to the film plane. The decrease in H/sub hfs/
indicates the decrease of the Fe particle size and the in-plane
orientation of H/sub hfs/ implies the non spherical shape of Fe
particles in the film. On the other hand, 400 keV Ar ion irradiation
induces the change from the superparamagnetic characteristics to the
ferromagnetic one; the ferromagnetic peaks showing the random
orientation of H/sub hfs/ appear and indicate the spherical growth of
Fe particles in Al/sub 2/O/sub 3/ matrix.

#e1#

#s1#
#abstract1116.txt#


Microprocessor architecture has entered the multicore era. Recently,
Hill and Marty presented a pessimistic view of multicore scalability.
Their analysis was based on Amdahl's law (i.e. fixed-workload
condition) and challenged readers to develop better models. In this
study, we analyze multicore scalability under fixed-time and
memory-bound conditions and from the data access (memory wall)
perspective. We use the same hardware cost model of multicore chips
used by Hill and Marty, but achieve very different and more optimistic
performance models. These models show that there is no inherent,
immovable upper bound on the scalability of multicore architectures.
These results complement existing studies and demonstrate that
multicore architectures are capable of extensive scalability. [All
rights reserved Elsevier].

#e1#

#s1#
#abstract1117.txt#


We present two improved results for scheduling batched parallel jobs on
multiprocessors with mean response time as the performance metric.
These results are obtained by using a generalized analysis framework
where the response time of the jobs is expressed in two contributing
factors that directly impact a scheduler's competitive ratio.
Specifically, we show that the scheduler IGDEQ is 3-competitive against
the optimal while AGDEQ is 5.24-competitive. These results improve the
known competitive ratios of 4 and 10, obtained by Deng et al. and by He
et al., respectively. For the common case where no fractional
allotments are allowed, we show that slightly larger competitive ratios
can be obtained by augmenting the schedulers with the round-robin
strategy. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1118.txt#


The paper reports in detail a methodology to fully exploit the
potential of a SIMD (Single Instruction Multiple Data) vector extension
in the evaluation of certain type of integrals, which occur in the
numerical solution of a Boundary Integral Equation (BIE) through the
Boundary Element Method (BEM). Specifically, we present an algorithm
for the fast evaluation of the integral coefficients appearing in the
assembly of the BEM system matrices, which represents an extremely
time-consuming task. The numerical scheme is tailored to the specific
structure of the integrals associated to a wave propagation phenomenon,
governed, in the time domain, by the D'Alembert equation. The reason of
this choice resides in the critical importance achieved by this class
of problems in many engineering applications. In particular, the
application framework this work belongs to is the design of
environmentally friendly commercial aircraft, for which the regulation
and certification restrictions are, nowadays, a key constraint
effecting even the conceptual phase of the design process. For the sake
of generality, we used here only the basic features of the SIMD vector
extension, common to all the specific architectures available on the
market. Particular attention is payed to the accuracy-related issues
arising from the use of the low-latency approximations of some of the
operators involved. The resulting algorithm minimizes the number of
operations involving operands belonging to the same register
("horizontal" or "intra-register" operations). Preliminary numerical
results reveal a remarkable speed-up of this highly-demanding part of
the solution process, close, in most of the cases, to the theoretical
peak. Standard multithreading techniques are additionally introduced to
further increase the performance on multiprocessors machines. [All
rights reserved Elsevier].

#e1#

#s1#
#abstract1119.txt#


In this contribution, high-throughput screening experiments are
reported to study the polymerization of different aromatic polyurethane
(PU) prepolymers. The prepared prepolymers were synthesized from
toluene diisocyanate (T80) with different molar mass polyether diols
and polyether triols, respectively. The reactions were performed in
solution using a Chemspeed Accelerator trade SLT106 automated parallel
synthesizer as well as in bulk to evaluate the high-throughput approach
for this kind of prepolymers. More than 100 samples were prepared and
characterized by GPC within 1 week labor time to investigate the
reaction kinetics and to compare the resulting trends obtained by
high-throughput experimentation (HTE) or by conventional, bulk
prepolymerization. The synthesis of the prepared prepolymers with a
linear (T80-Diol) or a branched (T80-Triol) structure followed a
second-order kinetic in solution but showed deviation from this
phenomenon in bulk under the selected reaction conditions, although the
same trends are observed in both cases. The calculation of the rate
constants allowed comparing the reactivity of different prepolymer
systems, which could have a significant influence on the industrial
application and processing of these materials. As a result, the HTE
approach was found to represent a powerful tool for the kinetic studies
of PU prepolymers. Moreover, in spite of the complexity of the curing
process, the results obtained by high-throughput solution
polymerization can be applied for evaluating the bulk polymerization. c
2009 Wiley Periodicals, Inc.

#e1#

#s1#
#abstract1121.txt#


This paper considers a production scheduling problem frequently found
in many industries whose raw materials are agricultural products. The
study focuses on a production system in processed canned fruit industry
as a case study. Common characteristics of the system include: (1) high
uncertainties of agricultural raw materials (fresh fruits), both in
terms of quality and quantity, which significantly affect the
production schedule that is usually planned in advance, (2) multiple
types of finished products (can sizes and fruit types), sharing the
same resources, thus makes their scheduling interdependent, and (3) the
shared resources (retorts) are non-identical which exist to provide
services to the bottleneck operation (sterilization) in parallel. The
paper has two interrelated objectives: to propose the use of real-time
scheduling methods based on dispatching rules for such systems, and to
demonstrate the use of computer simulation modeling to imitate the
actual production system and how to conduct computational experiment on
the simulation model to determine a set of appropriate dispatching
rules for the case study industry. Nine real-time, setup dependent,
dispatching rules are compared using two types of performance measures:
flow time and tardiness. The results show the effectiveness of this
approach to supporting the decision-making process in production
scheduling of canned fruit products. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1145.txt#


Combinatorial testing has been an active research area in recent years.
One challenge in this area is dealing with the combinatorial explosion
problem, which typically requires a very expensive computational
process to find a good test set that covers all the combinations for a
given interaction strength (f). Parallelization can be an effective
approach to manage this computational cost, that is, by taking
advantage of the recent advancement of multicore architectures. In line
with such alluring prospects, this paper presents a new deterministic
strategy, called multicore modified input parameter order (MC-MIPOG)
based on an earlier strategy, input parameter order generalized (IPOG).
Unlike its predecessor strategy, MCMIPOG adopts a novel approach by
removing control and data dependency to permit the harnessing of
multicore systems. Experiments are undertaken to demonstrate speedup
gain and to compare the proposed strategy with other strategies,
including IPOG The overall results demonstrate that MC-MIPOG
outperforms most existing strategies (IPOG, IPOF, IPOF2, IPOG-D, ITCH,
TConfig, Jenny, and TVG) in terms of test size within acceptable
execution time. Unlike most strategies, MC-MIPOG is also capable of
supporting high interaction strengths of t # 6.

#e1#

#s1#
#abstract1146.txt#


Fujitsu Laboratories of America has, over the course of many years,
worked to develop the frontier of binary decision diagram (BDD)
technology under a project called ParDD. Our technology allows us to
partition Boolean functions, represent them very compactly, and process
them on a massively parallel computing platform. It has been used to
create numerous applications in the field of electronic design
automation. Recently under this project we have developed a novel BDD
library where the storage requirement of each node closely tracks the
total size of the stored representation. The compact nature of this
data structure allows the solution of interesting problems to which
BDDs have seldom been applied before. For example, we have used our
library to create a compact inverted index, an essential matrix for
indexing documents in any corpus, including the World Wide Web. We have
also characterized its performance for Web query satisfaction in the
context of Web searches as well as for the creation of compact
representations of access control lists, a core component of Internet
routers.

#e1#

#s1#
#abstract1153.txt#


The paper deals with the following topics: digital simulation; grid
computing; production systems; neural networks; artificial
intelligence; soft computing; discrete-event systems; genetic
algorithms; operations research; bioinformatics; image processing;
speech processing; signal processing; knowledge management, enterprise
resource planning; social science; economic science; engineering
applications; power system application; transportation application;
virtual reality; data visualization; computer games; parallel
architectures; distributed architectures; Internet; ontologies;
electronic engineering applications; e-science; and e-systems.

#e1#

#s1#
#abstract1154.txt#


The following topics are dealt with: distributed real-time systems;
time-triggered systems; model-based development; dependable and secure
computing; timing analysis; component-based architectures;
configuration and adaptation; multi-core platforms; and validation and
verification.

#e1#

#s1#
#abstract1155.txt#


Parallel I/O is fast becoming a bottleneck to the research agendas of
many users of extreme scale parallel computers. The principle cause of
this is the concurrency explosion of high-end computation, coupled with
the complexity of providing parallel file systems that perform reliably
at such scales. More than just being a bottleneck, parallel I/O
performance at scale is notoriously variable, being influenced by
numerous factors inside and outside the application, thus making it
extremely difficult to isolate cause and effect for performance events.
In this paper, we propose a statistical approach to understanding I/O
performance that moves from the analysis of performance events to the
exploration of performance ensembles. Using this methodology, we
examine two I/O-intensive scientific computations from cosmology and
climate science, and demonstrate that our approach can identify
application and middleware performance deficiencies - resulting in more
than 4* run time improvement for both examined applications.

#e1#

#s1#
#abstract1156.txt#


Driven by the increasing demand for large-scale and high-performance
data protection, disk-based de-duplication storage has become a new
research focus of the storage industry and research community where
several new schemes have emerged recently. So far these systems are
mainly inline de-duplication approaches, which are centralized and do
not lend themselves easily to be extended to handle global
de-duplication in a distributed environment. We present DEBAR, a
de-duplication storage system designed to improve capacity, performance
and scalability for de-duplication backup/archiving. DEBAR performs
post-processing de-duplication, where backup streams are de-duplicated
and cached on server-disks through an in-memory preliminary filter in
phase I, and then completely de-duplicated in-batch in phase II. By
decentralizing fingerprint lookup and update, DEBAR supports a cluster
of servers to perform de-duplication backup in parallel, and is shown
to scale linearly in both write throughput and physical capacity,
achieving an aggregate throughput of 1.7GB/s and supporting a physical
capacity of 2PB with 16 backup servers.

#e1#

#s1#
#abstract1157.txt#


While Moore's law scaling continues to double transistor density every
technology generation, supply voltage reduction has essentially
stopped, increasing both power density and total energy consumed in
conventional microprocessors. Therefore, future processors will require
an architecture that can: a) take advantage of the massive amount of
transistors that will be available; and b) operate these transistors in
the near-threshold supply domain, thereby achieving near optimal
energy/computation by balancing the leakage and dynamic energy
consumption. Unfortunately, this optimality is typically achieved while
running at very low frequencies (i.e. 0:1 - 10 MHz) and with only one
computation executing per cycle, such that performance is limited.
Further, near-threshold designs suffer from severe process variability
that can introduce extremely large delay variations. In this paper, we
propose a near energy-optimal, stream processor family that relies on
massively parallel, near-threshold VLSI circuits and interconnect,
incorporating cooperative circuit/architecture techniques to tolerate
the expected large delay variations. Initial estimations from circuit
simulations show that it is possible to achieve greater than 1
Giga-Operations per second (1GOP/s) with less than 1 mW total power
consumption, enabling a new class of energy-constrained,
high-throughput computing applications.

#e1#

#s1#
#abstract1158.txt#


For the last few years, the major driving force behind the rapid
performance improvement of SSDs has been the increment of parallel bus
channels between a flash controller and flash memory packages inside
the solid-state drives (SSDs). However, there are other internal
parallelisms inside SSDs yet to be explored. In order to improve
performance further by utilizing the parallelism, this paper suggests
request rescheduling and dynamic write request mapping. Simulation
results with real workloads have shown that the suggested schemes
improve the performance of the SSDs by up to 15% without any additional
hardware support.

#e1#

#s1#
#abstract1159.txt#


Flash memory solid-state disks (SSDs) are replacing hard disk drives
(HDDs) in mobile computing systems because of their lower power
consumption, faster random access, and greater shock resistance. We
describe Hydra, a high-performance flash memory SSD architecture that
translates the parallelism inherent in multiple flash memory chips into
improved performance, by means of both bus-level and chip-level
interleaving. Hydra has a prioritized structure of memory controllers,
consisting of a single high-priority foreground unit, to deal with read
requests, and multiple background units, all capable of autonomous
execution of sequences of high-level flash memory operations. Hydra
also employs an aggressive write buffering mechanism based on block
mapping to ensure that multiple flash memory chips are used
effectively, and also to expedite the processing of write requests.
Performance evaluation of an FPGA implementation of the Hydra SSD
architecture shows that its performance is more than 80 percent better
than the best of the comparable HDDs and SSDs that we considered.

#e1#

#s1#
#abstract1160.txt#


Multi-band orthogonal frequency-division multiplexing (MB-OFDM)
ultra-wideband (UWB) technology offers large throughput, low latency
and has been adopted in wireless audio/video (AV) network products. The
complexity and power consumption, however, are still major hurdles for
the technology to be widely adopted. In this paper, we propose a
unified synchronizer design targeted for MB-OFDM transceiver that
achieves high performance with low implementation complexity. The key
component of the proposed synchronizer is a /i parallel/
auto-correlator structure in which multiple ACF units are instantiated
and their outputs are /i shared/ by functional blocks in the
synchronizer, including preamble signal detection, time-frequency code
identification, symbol timing, carrier frequency offset estimation and
frame synchronization. This common structure not only reduces the
hardware cost but also minimizes the number of operations in the
functional blocks in the synchronizer as the results of a large portion
of computation can be shared among different functional blocks. To
mitigate the effect of narrowband interference (NBI) on UWB systems, we
also propose a low-complexity ACF-based frequency detector to
facilitate the design of (adaptive) notch filter in analog/digital
domain. The theoretical analysis and simulation show that the
performance of the proposed design is close to optimal, while the
complexity is significantly reduced compared to existing work.

#e1#

#s1#
#abstract1161.txt#


The estimation of motion and structure from stereo-video streams is
revisited for applications in ocean exploration and seafloo Gamma
mapping. Operational constraints require real-time processing to enable
adaptive trajectory planning, and robust estimation to navigate optimal
paths for collection of useful underwater stereo data. While the
traditional joint estimation of motion and structure by way of an
extended Kalman filter (EKF) provides a suitable recursive framework,
the cubic computation growth with the number of feature tracks is a
serious bottleneck. By treating the motion and structure as the states
of two coupled filters with stereo feature correspondences as
observations, a dual estimator is devised with the performance of joint
estimation and computational complexity proportional to the number of
features. We favor a sequential implementation to ensure unbiased
estimation, in contrast to two parallel dual estimation schemes that
generally produce biased updates. Stochastic stability can be
established in terms of conditions on initial estimation error, bound
on observation noise covariance, observation nonlinearity, and modeling
error. Moreover, dynamic features can be treated effectively and
efficiently by the removal or addition to a bank of filters, one
assigned per feature. Experimental results with synthetic and several
real data sets are presented to demonstrate the merits of the proposed
recursive dual EKF-based estimator.

#e1#

#s1#
#abstract1162.txt#


A simple, low-cost instrument that measures impedance and phase angle
was used along with a parallel-plate capacitance system to estimate the
moisture content (MC) of in-shell peanuts and yellow-dent field corn.
Moisture content of the field crops is important and is measured at
various stages of their processing and storage. A sample of about 150 g
of in-shell peanuts or corn was placed separately between a set of
parallel plate electrodes and the impedance and phase angle of the
system were measured at frequencies 1 and 5 MHz. A semi-empirical
equation was developed for peanuts and corn separately using the
measured impedance and phase angle values, and the computed capacitance
and the MC values obtained by standard air-oven method. The multilinear
regression (MLR) method was used for the empirical equation development
using an Unscrambler 9.7 data analyzer. In this paper, a low-cost
impedance analyzer designed and assembled in our laboratory was used to
measure the impedance and phase angles. MC values of corn samples in
the moisture range of 7% to 18% and in-shell peanuts in the moisture
range of 9% to 20%, not used in the calibration, were predicted by the
equations and compared with their standard air-oven values. For over
96% of the samples tested from both crops, the predicted MC values were
within 1% of the air-oven values. This method, being nondestructive and
rapid, will have considerable application in the drying and storage
processes for peanuts, corn, and similar field crops.

#e1#

#s1#
#abstract1163.txt#


This paper presents a video metrology approach using an uncalibrated
single camera that is either stationary or in planar motion. Although
theoretically simple, measuring the length of even a line segment in a
given video is often a difficult problem. Most existing techniques for
this task are extensions of single image-based techniques and do not
achieve the desired accuracy especially in noisy environments. In
contrast, the proposed algorithm moves line segments on the reference
plane to share a common endpoint using the vanishing line information
followed by fitting multiple concentric circles on the image plane. A
fully automated real-time system based on this algorithm has been
developed to measure vehicle wheelbases using an uncalibrated
stationary camera. The system estimates the vanishing line using
invariant lengths on the reference plane from multiple frames rather
than the given parallel lines, which may not exist in videos. It is
further extended to a camera undergoing a planar motion by
automatically selecting frames with similar vanishing lines from the
video. Experimental results show that the measurement results are
accurate enough to classify moving vehicles based on their size.

#e1#

#s1#
#abstract1164.txt#


This book presents an overview of the design and implementation of
transactional memory systems. The topics include:transactional memory;
parallel programming; concurrent programming; compilers; programming
languages; computer architecture; computer hardware; nonblocking
algorithms; lock-free data structures; cache coherence; and
synchronization.

#e1#

#s1#
#abstract1165.txt#


The book provides a comprehensive description of the specific problems
arising in cross-language information retrieval. The topics
include:multilingual information retrieval; query translation; document
translation; translation model; machine translation; statistical
machine translation; dictionary-based translation; parallel corpus;
comparable corpus; query expansion; transliteration; and mining of
translation resources.

#e1#

#s1#
#abstract1167.txt#


The following topics are dealt with: performance modelling; routing
mechanism; distributed systems; parallel systems; service oriented
architecture; communication technology; protocols; autonomic network
systems; Internet computing; ad hoc networks; wireless sensor networks;
privacy management; mobile networks; multimedia systems; pervasive
computing; ubiquitous computing; security intrusion detection; Web
services; security authentication; access control; WLAN; OFDM systems;
grid computing; P2P computing; database processing; data mining; and
semantics collaborative systems.

#e1#

#s1#
#abstract1168.txt#


The following topics are dealt with: FPGA; computer vision; computer
graphics processing; run-time systems; supercomputing; open-source
tools; CAD tools; computer application development; machine learning;
string matching; networking; computer architectures and encryption.

#e1#

#s1#
#abstract1189.txt#


Biological systems that are capable of performing computational
operations could be of use in bioengineering and nanomedicine, and DNA
and other biomolecules have already been used as active components in
biocomputational circuits. There have also been demonstrations of
DNA/RNA-enzyme-based automatons, logic control of gene expression, and
RNA systems for processing of intracellular information. However, for
biocomputational circuits to be useful for applications it will be
necessary to develop a library of computing elements, to demonstrate
the modular coupling of these elements, and to demonstrate that this
approach is scalable. Here, we report the construction of a DNA-based
computational platform that uses a library of catalytic nucleic acids
(DNAzymes), and their substrates, for the input-guided dynamic assembly
of a universal set of logic gates and a half-adder/half-subtractor
system. We demonstrate multilayered gate cascades, fan-out gates and
parallel logic gate operations. In response to input markers, the
system can regulate the controlled expression of anti-sense molecules,
or aptamers, that act as inhibitors for enzymes.

#e1#

#s1#
#abstract1195.txt#


Distributed and multiprocessor computing is a a basic need for
increasingly complex computational requirements. Scheduling the tasks
efficiently is as important in multiprocessor systems as getting the
correct results. In this paper, we present a heuristic algorithm for
optimized parallelization of tasks in multiprocessor environments. This
heuristic can be applied in such multiprocessor environments where task
and resource information is completely know at time of scheduling or
incomplete task information is known at scheduling time. For the case
of incomplete task information, we have given a method of using static
cost analysis to get different sets of parallel tasks and one set among
these alternatives is chosen in runtime based on the value of input
parameter. We have verified our approach using a prolog (CLP) program
for random sets of tasks and resource with 100% positive results. Our
heuristic algorithm is quite simple yet effective which makes it
different than already existing scheduling algorithms, heuristics and
genetic algorithms.

#e1#

#s1#
#abstract1196.txt#


The represented article considers the models of optimization of
operation processes, namely, the problem of defining an optimal
sequence of accepted orders in the cases of the single and multiple
parallel working-places. In the article there are considered
economical-mathematical models and corresponding algorithms to solve a
given task.

#e1#

#s1#
#abstract1197.txt#


The paper deals with the following topics: supervised learning;
unsupervised learning; neurobiology; neurosciences; neuro-fuzzy
systems; Takagi-Sugeno models; information systems; image processing;
parallel computing applications; control applications; financial
mathematics; industrial measurement; large scale systems; quintitative
methods; robotics; mechatronics; multiobjective programming; game
theory; electromagnetics; and risk management.

#e1#

#s1#
#abstract1205.txt#


To measure the thermal emission from stratospheric minor species with
high sensitivity, the Superconducting Submillimeter-wave Limb-Emission
Sounder (SMILES) aboard the Japanese Experiment Module (JEM) of the
International Space Station (ISS) carries 4 K cooled
Superconductor-Insulator-Superconductor (SIS) mixers. The major feature
of the SMILES is its high-sensitive measurement ability with low system
noise temperature less than 700 K.As a part of the ground system for
the SMILES, a level 2 data processing system (DPS-L2) has been
developed. It retrieves the density distributions of the target species
from calibrated spectra in near-real-time. The retrieval process
consists of two parts: the forward model, which computes radiative
transfer, and the inverse model, which deduces atmospheric states.
Since the forward model must provide the most accurate basis for
results and be implemented under limited computing resources, the
forward model algorithm for an operational system has to be accurate
and fast. Hence, the algorithm is improved (1) by designing accurate
instrument functions such as the instrumental field of view (FOV),
sideband rejection ratio of sideband separator, and spectral responses
of acousto-optic spectrometer (AOS) and (2) by optimizing radiative
transfer calculation.This paper presents the development of the DPS-L2
along with the details on its algorithm and the algorithm performance.
The accuracy of this algorithm is better than 1%, and the processing
time for single-scan spectra is less than 1 min with eight parallel
processings using a 3.16-GHz Quad-Core Intel Xeon processor. Thus, this
algorithm is suitable for the SMILES measurement. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract1217.txt#


Due to its high-computational complexity and poor parallelism, the
context-based adaptive binary arithmetic coder (CABAC) increasingly
poses a bottleneck in the large-scale parallel video encoder like H.264
on a manycore platform. Motivated by the organization of the contexts
models in CABAC, this letter presents a tri-thread parallel evolution
of CABAC. The evolutional coder, which is named P3-CABAC, statically
divides syntax elements into three predefined groups, each of which
forms a parallel thread. Since the P3-CABAC is thread-level
parallelizable and implementation-friendly for manycore processors, it
presents a rational consideration for entropy coding in block-based
video encoders in the manycore era.

#e1#

#s1#
#abstract1218.txt#


Parallel-to-cascading optical code (OC) label processing based on fiber
Bragg grating is proposed and experimentally demonstrated. With this
technique, the length of the multibit OC label is the same as that of
the one-bit OC label; only one encoder and one decoder are needed for
the multibit OC label's generation and recognition; the relatively
low-speed recognition patterns also reduce the complexity of generating
the control signal. The experimental results show that both the ratio
of autocorrelation peak to cross-correlation peak (ACP/CCP) and the
ratio of ACP to maximum autocorrelation wing are higher than 8 dB in
the multibit label's decoded signal, so the parallel-to-cascading OC
label processing is practical for optical packet switching networks.

#e1#

#s1#
#abstract1225.txt#


Background: RNA structure prediction problem is a computationally
complex task, especially with pseudo-knots. The problem is well-studied
in existing literature and predominantly uses highly coupled Dynamic
Programming (DP) solutions. The problem scale and complexity become
embarrassingly humungous to handle as sequence size increases. This
makes the case for parallelization. Parallelization can be achieved by
way of networked platforms (clusters, grids, etc) as well as using
modern day multi-core chips. Methods: In this paper, we exploit the
parallelism capabilities of the IBM Cell Broadband Engine to
parallelize an existing Dynamic Programming (DP) algorithm for RNA
secondary structure prediction. We design three different
implementation strategies that exploit the inherent data, code and/or
hybrid parallelism, referred to as C-Par, D-Par and H-Par, and analyze
their performances. Our approach attempts to introduce parallelism in
critical sections of the algorithm. We ran our experiments on SONY Play
Station 3 (PS3), which is based on the IBM Cell chip. Results: Our
results suggest that introducing parallelism in DP algorithm allows it
to easily handle longer sequences which otherwise would consume a large
amount of time in single core computers. The results further
demonstrate the speed-up gain achieved in exploiting the inherent
parallelism in the problem and also elicits the advantages of using
multi-core platforms towards designing more sophisticated methodologies
for handling a fairly long sequence of RNA. Conclusion: The speed-up
performance reported here is promising, especially when sequence length
is long. To the best of our literature survey, the work reported in
this paper is probably the first-of-its-kind to utilize the IBM Cell
Broadband Engine (a heterogeneous multi-core chip) to implement a DP.
The results also encourage using multi-core platforms towards designing
more sophisticated methodologies for handling a fairly long sequence of
RNA to predict its secondary structure.

#e1#

#s1#
#abstract1226.txt#


To improve detecting rates and reduce false detection of distributed
network intrusion detection system, and to improve parallel processing
ability of distributed intrusion detection system, co-evaluation
computation-based distributed intrusion detection system is proposed.
Optimized immune detecting method is used to reduce redundancy of
detector. Multi-agents evaluation computation is used to enhance
self-learning ability and self-adaptation ability of network intrusion
detection. Co-evaluation technology is used to speed co-evaluating of
multi-agents in network intrusion detection system and then improve
evaluating ability of distributed system. Experiments verify the
validity of the method.

#e1#

#s1#
#abstract1227.txt#


Computational material design requires efficient algorithms and
high-speed computers for calculating and predicting material
properties. The orbital-free first principles calculation (OF-FPC)
method, which is a tool for calculating and designing material
properties, is an O(N) method and is suitable for large-scaled systems.
The stagnation in the development of CPU devices with high mobility of
electron carriers has driven the development of parallel computing and
the production of CPU devices with finer spaced wiring. We, for the
first time, propose another method to accelerate the computation using
Graphics Processing Unit (GPU). The implementation of the Fast Fourier
Transform (CUFFT) library that uses GPU, into our in-house OF-FPC code,
reduces the computation time to half of that of the CPU.

#e1#

#s1#
#abstract1228.txt#


Rapid development of infrared detector arrays caused a need to develop
robust signal processing chain able to perform operations on infrared
image in real-time. Every infrared detector array suffers from
so-called nonuniformity, which has to be digitally compensated by the
internal circuits of the camera. Digital circuit also has to detect and
replace signal from damaged detectors. At the end the image has to be
prepared for display on external display unit. For the best comfort of
viewing the delay between registering the infrared image and displaying
it should be as short as possible. That is why the image processing has
to be done with minimum latency. This demand enforces to use special
processing techniques like pipelining and parallel processing. Designed
infrared processing module is able to perform standard operations on
infrared image with very low latency. Additionally modular design and
defined data bus allows easy expansion of the signal processing chain.
Presented image processing module was used in two camera designs based
on uncooled microbolometric detector array form ULIS and cooled photon
detector from Sofradir. The image processing module was implemented in
FPGA structure and worked with external ARM processor for control and
coprocessing. The paper describes the design of the processing unit,
results of image processing, and parameters of module like power
consumption and hardware utilization.

#e1#

#s1#
#abstract1229.txt#


We re-examine the problem of load balancing in conservatively
synchronized parallel, discrete- event simulations executed on
high-performance computing clusters, focusing on simulations where
computational and messaging load tend to be spatially clustered. Such
domains are frequently characterized by the presence of geographic
"hot-spots'' - regions that generate significantly more simulation
events than others. Examples of such domains include simulation of
urban regions, transportation networks and networks where interaction
between entities is often constrained by physical proximity. Noting
that in conservatively synchronized parallel simulations, the speed of
execution of the simulation is determined by the slowest ( i.e most
heavily loaded) simulation process, we study different partitioning
strategies in achieving equitable processor-load distribution in
domains with spatially clustered load. In particular, we study the
effectiveness of partitioning via spatial scattering to achieve optimal
load balance. In this partitioning technique, nearby entities are
explicitly assigned to different processors, thereby scattering the
load across the cluster. This is motivated by two observations, namely,
(i) since load is spatially clustered, spatial scattering should,
intuitively, spread the load across the compute cluster, and (ii) in
parallel simulations, equitable distribution of CPU load is a greater
determinant of execution speed than message passing overhead. Through
large-scale simulation experiments - both of abstracted and real
simulation models - on high performance clusters, we observe that
scatter partitioning - even with its greatly increased messaging
overhead - often significantly outperforms more conventional spatial
partitioning techniques that seek to reduce messaging overhead.
Further, even if hot-spots change over the course of the simulation, if
the underlying feature of spatial clustering is retained, load
continues to be balanced with spatial scattering leading us to the
observation that spatial scattering can often obviate the need for
dynamic load balancing.

#e1#

#s1#
#abstract1230.txt#


Stream processing is an important emerging computational model for
performing complex operations on and across multi-source, high volume,
unpredictable dataflows. We present Flow, a platform for parallel and
distributed stream processing system simulation that provides a
flexible modeling environment for analyzing stream processing
applications. The Flow stream processing system simulator is a high
performance, scalable simulator that automatically parallelizes chunks
of the model space and incurs near zero synchronization overhead for
stream application graphs that exhibit feed-forward behavior. We show
promising multi-threaded and multi-process event rates exceeding 80
million events per second on a cluster with 256 processor cores.

#e1#

#s1#
#abstract1231.txt#


The spatial scale, runtime speed and behavioral detail of epidemic
outbreak simulations together require the use of large-scale parallel
processing. In this paper, an optimistic parallel discrete event
execution of a reaction-diffusion simulation model of epidemic
outbreaks is presented, with an implementation over the mu sik
simulator. Rollback support is achieved with the development of a novel
reversible model that combines reverse computation with a small amount
of incremental state saving. Parallel speedup and other runtime
performance metrics of the simulation are tested on a small
(8,192-core) Blue Gene/P system, while scalability is demonstrated on
65,536 cores of a large Cray XT5 system. Scenarios representing large
population sizes (up to several hundred million individuals in the
largest case) are exercised.

#e1#

#s1#
#abstract1232.txt#


Multi-core processors are commonly available now, but most traditional
computer architectural simulators still use single-thread execution. In
this paper we use parallel discrete event simulation (PDES) to speedup
a cycle-accurate event-driven many-core processor simulator. Evaluation
against the sequential version shows that the parallelized one achieves
an average speedup of 10.9x (up to 13.6x) running SPLASH-2 kernel on a
16-core host machine, with cycle counter differences of less than 0.1%.
Moreover, super-linear speedups are achieved between running 1 thread
and 8 threads due to reduced overhead of insert-event-to-queue time and
increased cache size in parallel processing. We conclude that PDES
could be an attractive option for achieving fast cycle-accurate
many-core processor simulations.

#e1#

#s1#
#abstract1233.txt#


Katsevich proposed the first theoretically exact reconstruction formula
for spiral cone-beam CT which is computation quite intensive. The main
bottleneck is that the memory storing the projection data has to be
accessed for many times. Given that one single source position s
belongs to S different pi -intervals, and the corresponding pi -line
contains N voxels, then the projection data at position s has to be
accessed for O(S*N) times. In this work, we propose a parallel
algorithm which can reduce the times accessing the memory down to one.

#e1#

#s1#
#abstract1234.txt#


A simple parallel ray approximation based stochastic channel model for
multiple-input multiple-output (MIMO) ultra-wideband (UWB) system is
proposed. Considering the attenuation of the multipath amplitudes, the
phase shift and the difference of arrival time, the IEEE 802.15.3a
standard channel model for single-input single-out (SISO) scenario is
extended to the spatially correlated MIMO channel. The attenuation of
the multipath amplitudes and the arrival time differences are
determined by the structure of spatial linear arrays, the angle of
departure (AOD) and the angle of arrival (AOA) of multipath components.
The AOD and AOA can be approximated by Laplacian distribution according
to the measurement results. In order to verify the channel model, the
spatial correlation, the MIMO channel capacity, the bite-error-rate
(BER) performance of spatial multiplexing (SM) UWB system and the BER
performance of space time block coding (STBC) UWB system under the
proposed channel model are compared with those under the measured
channels. The measurement of MIMO UWB channel is carried out in a lab
environment with virtual antennas. The results show that the simple
MIMO UWB channel model matches the measured channel very well and
performs better than the correlation-based channel model (CBCM).

#e1#

#s1#
#abstract1235.txt#


Sorting is a kernel algorithm for a wide range of applications. In this
paper, we present a new algorithm, GPU-Warpsort, to perform
comparison-based parallel sort on Graphics Processing Units (GPUs). It
mainly consists of a bitonic sort followed by a merge sort. Our
algorithm achieves high performance by efficiently mapping the sorting
tasks to GPU architectures. Firstly, we take advantage of the
synchronous execution of threads in a warp to eliminate the barriers in
bitonic sorting network. We also provide sufficient homogeneous
parallel operations for all the threads within a warp to avoid branch
divergence. Furthermore, we implement the merge sort efficiently by
assigning each warp independent pairs of sequences to be merged and by
exploiting totally coalesced global memory accesses to eliminate the
bandwidth bottleneck. Our experimental results indicate that
GPU-Warpsort works well on different kinds of input distributions, and
it achieves up to 30% higher performance than previous optimized
comparison-based GPU sorting algorithm on input sequences with millions
of elements.

#e1#

#s1#
#abstract1236.txt#


Uintah is a highly parallel and adaptive multi-physics framework
created by the Center for Simulation of Accidental Fires and Explosions
in Utah. Uintah, which is built upon the Common Component Architecture,
has facilitated the simulation of a wide variety of fluid-structure
interaction problems using both adaptive structured meshes for the
fluid and particles to model solids. Uintah was originally designed
for, and has performed well on, about a thousand processors. The
evolution of Uintah to use tens of thousands processors has required
improvements in memory usage, data structure design, load balancing
algorithms and cost estimation in order to improve strong and weak
scalability up to 98,304 cores for situations in which the mesh used
varies adaptively and also cases in which particles that represent the
solids move from mesh cell to mesh cell.

#e1#

#s1#
#abstract1237.txt#


Graphics Processing Units (GPUs) are massively parallel, many-core
processors with tremendous computational power and very high memory
bandwidth. With the advent of general purpose programming models such
as NVIDIA's CUDA and the new standard OpenCL, general purpose
programming using GPUs (GPGPU) has become very popular. However, the
GPU architecture and programming model have brought along with it many
new challenges and opportunities for compiler optimizations. One such
classical optimization is loop unrolling. Current GPU compilers perform
limited loop unrolling. In this paper, we attempt to understand the
impact of loop unrolling on GPGPU programs. We develop a
semi-automatic, compile-time approach for identifying optimal unroll
factors for suitable loops in GPGPU programs. In addition, we propose
techniques for reducing the number of unroll factors evaluated, based
on the characteristics of the program being compiled and the device
being compiled to. We use these techniques to evaluate the effect of
loop unrolling on a range of GPGPU programs and show that we correctly
identify the optimal unroll factors. The optimized versions run up to
70 percent faster than the unoptimized versions.

#e1#

#s1#
#abstract1238.txt#


The phaser construct is a unification of collective and point-to-point
synchronization with dynamic parallelism. This construct gives each
task the option of synchronizing on a phaser in signal-only/wait-only
mode for producer/consumer synchronization or signal-wait mode for
barrier synchronization. A phaser accumulator is a reduction construct
that works with phasers in a phased setting. Phasers and accumulators
support dynamic parallelism i.e., they allow dynamic addition and
removal of tasks from the synchronizations and reductions that they
support. Past implementations of phasers and phaser accumulators have
used a single master task to advance a phaser to the next phase and to
perform computations for lazy reductions, while also supporting dynamic
parallelism. Though the single master approach provides an effective
solution for modest levels of parallelism, it quickly becomes a
scalability bottleneck as the number of threads increases. To address
this limitation, we propose an approach based on hierarchical phasers
for scalable synchronization and hierarchical accumulators for scalable
reduction. Our approach also includes tunable initialization parameters
that specify the degree and number of tiers for the phaser hierarchy,
thereby allowing different values to be chosen for different platforms.
Our performance results show significant scalability benefits from our
approach. To the best of our knowledge, this is the first approach to
support hierarchical synchronization and reductions in the presence of
dynamic parallelism.

#e1#

#s1#
#abstract1239.txt#


The Coarse-Grained Monte Carlo (CGMC) method is a multi-scale
stochastic mathematical and simulation framework for spatially
distributed systems. CGMC simulations are important tools for studying
phenomena such as catalysis, crystal growth, surface diffusion, phase
transitions on single crystals, and cell membrane receptor dynamics. In
parallel CGMC, the tau-leap method is used for parallel simulations
that are executed on traditional CPU clusters in a master-slave
setting. Unfortunately the communications between master and slaves
negatively impact speedup and scalability. In this paper, we explore
the potentials of GPUs for the tau-leap method and we present an
extensive performance evaluation that leads to the most suitable degree
of parallelism for this method under different simulation profiles. We
show how the efficient parallelization of the tau-leap method for GPUs
includes (1) the redefinition of its data structures, (2) the redesign
of its algorithm, and (3) the selection of the most appropriate degree
of parallelism (i.e., fine-grained or course-gained) on a single GPU or
multiple GPUs. Exceptional performance improvements can thus be
achieved for this method.

#e1#

#s1#
#abstract1240.txt#


System noise or Jitter is the activity of hardware, firmware, operating
system, runtime system, and management software events. It is shown to
disproportionately impact application performance in current generation
large-scale clustered systems running general-purpose operating systems
(GPOS). Jitter mitigation techniques such as co-scheduling jitter
events across operating systems improve application performance but
their effectiveness on future petascale systems is unknown. To
understand if existing co-scheduling solutions enable scalable
petascale performance, we construct two complementary jitter models
based on detailed analysis of system noise from the nodes of a
large-scale system running a GPOS. We validate these two models using
experimental data from a system consisting of 128 GPOS instances with
4096 CPUs. Based on our models, we project a minimum slowdown of 2.1%,
5.9%, and 11.5% for applications executing on a similar one petaflop
system running 1024 GPOS instances and having global synchronization
operations once every 1000 msec, 100 msec, and 10 msec, respectively.
Our projections indicate that additional system noise mitigation
techniques are required to contain the impact of jitter on
multi-petaflop systems, especially for tightly synchronized
applications.

#e1#

#s1#
#abstract1241.txt#


Multi-pattern string matching remains a major performance bottleneck in
network intrusion detection and anti-virus systems for high-speed deep
packet inspection (DPI). Although Aho-Corasick deterministic finite
automaton (AC-DFA) based solutions produce deterministic throughput and
are widely used in today's DPI systems such as Snort [1] and ClamAV
[2], the high memory requirement of AC-DFA (due to the large number of
state transitions in AC-DFA) inhibits efficient hardware implementation
to achieve high performance. Some recent work [3], [4] has shown that
the AC-DFA can be reduced to a character trie that contains only the
forward transitions by incorporating pipelined processing. But they
have limitations in either handling long patterns or extensions to
support multi-character input per clock cycle to achieve high
throughput. This paper generalizes the problem and proves formally that
a linear pipeline with H stages can remove all cross transitions to the
top H levels of a AC-DFA. A novel and scalable pipeline architecture
for memory-efficient multi-pattern string matching is then presented.
The architecture can be easily extended to support multi-character
input per clock cycle by mapping a compressed AC-DFA [5] onto multiple
pipelines. Simulation using Snort and ClamAV pattern sets shows that a
8-stage pipeline can remove more than 99% of the transitions in the
original AC-DFA. The implementation on a state-of-the-art field
programmable gate array (FPGA) shows that our architecture can store on
a single FPGA device the full set of string patterns from the latest
Snort rule set. Our FPGA implementation sustains 10+ Gbps throughput,
while consuming a small amount of on-chip logic resources. Also
desirable scalability is achieved: the increase in resource requirement
of our solution is sub-linear with the throughput improvement.

#e1#

#s1#
#abstract1242.txt#


Transactional Memory (TM) has attracted considerable attention because
it promises to increase programmer productivity by making it easier to
write correct parallel programs. To maintain correctness in the face of
concurrency, detecting conflicts among simultaneously running
transactions is an essential element. Hardware signatures have been
proposed as an area-efficient mechanism for conflict detection. A
signature can summarize an unbounded amount of addresses and misses no
conflicts, but could falsely declare conflicts even when no true
conflict exists (false positives) due to aliasing and occupancy.
Previous signature designs assume that false positives are destructive
to performance and attempt to reduce the total number of false
positives. In this paper, we show that some false positives can be
helpful to performance by triggering the early abortion of a
transaction which would encounter a true conflict later anyway. Based
on this observation, we propose an adaptive grain signature to improve
performance by dynamically changing the range of address keys based on
the history. With the use of adaptive grain signatures, we can increase
the number of performance-friendly false positives as well as decrease
the number of performance-destructive false positives.

#e1#

#s1#
#abstract1243.txt#


Scheduling of large-scale, distributed topology-aware applications
requires that not only the properties of the requested machines be
considered, but also the properties of the machines' interconnections.
This requirement severely complicates the scheduling process, as even a
matching between a single multi-processors task and available machines
in a single time slot becomes an NP-complete problem with no polynomial
approximation. In this paper we propose a complete scheduling framework
for multi-cluster, heterogeneous environments that provides, in
practice, an efficient solution for the scheduling of topology-aware
applications. The proposed framework is very flexible as it is composed
of pluggable components and can be easily configured to support a
variety of scheduling policies. W e also describe three novel
scheduling and coallocation algorithms that were developed and plugged
into the framework. The proposed scheduling framework was integrated
into the QosCosGrid system, where it is used as the main
decision-making module.

#e1#

#s1#
#abstract1244.txt#


Multi-core organizations increasingly support multiple threads per
core. Threads on a core usually share a single first-level data cache,
so thread schedulers must try to minimize cache contention among
threads. While this has been studied for concurrent threads with
disjoint working sets, the problem has not been addressed for
multi-threaded data-parallel workloads in which threads can be
scheduled or constructed to improve inter-thread cache sharing. This
paper proposes the symbiotic affinity scheduling (SAS) algorithm in
which work is first partitioned according to the number of cores (i.e.,
the number of caches), and these partitions are then subdivided and
scheduled among each core's available thread contexts so that threads
sharing a core operate on neighboring elements to maximize cache
locality. We demonstrate this concept with a series of data-parallel
benchmarks. Simulations on M5 achieve an average speedup of 1.69* and
36% energy savings over conventional scheduling techniques that are
oblivious to whether threads share a cache. Even compared to an
approach that extends oblivious scheduling to ensure that the sum of
the threads' working sets fits in the cache, symbiotic affinity
scheduling is able to exploit greater temporal locality and provide 30%
performance gains on average. Symbiosis also outperforms adaptive
contention reduction techniques by 17%.

#e1#

#s1#
#abstract1245.txt#


Exploiting parallelism in route planning algorithms is a challenging
algorithmic problem with obvious applications in mobile navigation and
timetable information systems. In this work, we present a novel
algorithm for the so-called one-to-all profile-search problem in public
transportation networks. It answers the question for all fastest
connections between a given station S and any other station at any time
of the day in a single query. This algorithm allows for a very natural
parallelization, yielding excellent speed-ups on standard multi-core
servers. Our approach exploits the facts that first, time-dependent
travel-time functions in such networks can be represented as a special
class of piecewise linear functions, and that second, only few
connections from S are useful to travel far away. Introducing the
connection-setting property, we are able to extend DIJKSTRA's algorithm
in a sound manner. Furthermore, we also accelerate station-tostation
queries by preprocessing important connections within the public
transportation network. As a result, we are able to compute all
relevant connections between two random stations in a complete public
transportation network of a big city (Los Angeles) on a standard
multi-core server in less than 55 ms on average.

#e1#

#s1#
#abstract1246.txt#


Energy efficiency and parallel I/O performance have become two critical
measures in high performance computing (HPC). However, there is little
empirical data that characterize the energy-performance behaviors of
parallel I/O workload. In this paper, we present a methodology to
profile the performance, energy, and energy efficiency of parallel I/O
access patterns and report our findings on the impacting factors of
parallel I/O energy efficiency. Our study shows that choosing the right
buffer size can change the energy-performance efficiency by up to 30
times. High spatial and temporal spacing can also lead to significant
improvement in energy-performance efficiency (about 2X). We observe CPU
frequency has a more complex impact, depending on the IO operations,
spatial and temporal, and memory buffer size. The presented methodology
and findings are useful for evaluating the energy efficiency of I/O
intensive applications and for providing a guideline to develop energy
efficient parallel I/O technology.

#e1#

#s1#
#abstract1247.txt#


Next-generation high throughput sequencing instruments are capable of
generating hundreds of millions of reads in a single run. Mapping those
reads to a reference genome is an extremely compute-intensive process
that takes more than a day on a modern computer even when the accuracy
of the results is traded off to speed up the execution. In this work,
we explore various data distribution strategies for parallel execution
of three state-of-the-art mapping tools, namely Bowtie, BWA and SOAP2,
that are based on the Burrows-Wheeler Transformation. We report on the
performance of these strategies and show that the best strategy depends
on the input scenario as well as the relative efficiency of the tools
in the indexing and matching steps of the mapping process. The
parallelization strategies investigated in this paper are general and
can easily be applied to different mapping algorithms. With the
availability of parallel execution methods, it will be possible to
carry out more intensive computations that cannot be accomplished in a
reasonable time using sequential tools, including mapping with larger
mismatch tolerance.

#e1#

#s1#
#abstract1248.txt#


With the development of high-performance computing, I/O issues have
become the bottleneck for many massively parallel applications. This
paper investigates scalable parallel I/O alternatives for massively
parallel partitioned solver systems. Typically such systems have
synchronized "loops" and will write data in a well defined block I/O
format consisting of a header and data portion. Our target use for such
an parallel I/O subsystem is checkpoint-restart where writing is by far
the most common operation and reading typically only happens during
either initialization or during a restart operation because of a system
failure. We compare four parallel I/O strategies: 1 POSIX File Per
Processor (1PFPP), a synchronized parallel IO library (syncIO),
"Poor-Man's" Parallel I/O (PMPIO) and a new "reduced blocking" strategy
(rbIO). Performance tests using real CFD solver data from PHASTA (an
unstructured grid finite element Navier-Stokes solver) show that the
syncIO strategy can achieve a read bandwidth of 6.6GB/Sec on Blue
Gene/L using 16K processors which is significantly faster than 1PFPP or
PMPIO approaches. The serial "token-passing" approach of PMPIO yields a
900 MB/sec write bandwidth on 16K processors using 1024 files and 1PFPP
achieves 600 MB/sec on 8K processors while the "reduced-blocked" rbIO
strategy achieves an actual writing performance of 2.3GB/sec and
perceived/latency hiding writing performance of more than 21,000 GB/sec
(i.e., 21TB/sec) on a 32,768 processor Blue Gene/L.

#e1#

#s1#
#abstract1249.txt#


Because of the very favorable price to performance ratio of the GPUs, a
popular parallel programming configuration today is a cluster of GPUs.
However, extracting performance on such a configuration would typically
require programming in both MPI and CUDA, thus requiring a high degree
of expertise and effort. It is clearly desirable to be able to support
higher-level programming of this emerging high-performance computing
platform. This paper reports on a code generation system that can
translate data mining applications on a GPU cluster. Our work is driven
by the observation that a common processing structure, that of
generalized reductions, fits a large number of popular data mining
algorithms. In our solution, the programmers simply need to specify the
sequential reduction loop(s) with some additional information about the
parameters. We use program analysis and code generation to
automatically map the applications to the API of FREERIDE, which is a
middleware for parallel data mining. We also automatically generate
CUDA code for using the GPU on each node of the cluster. We have
evaluated our system using two popular data mining applications,
k-means clustering and Principal Component Analysis (PCA). We observed
good scalability over the number of computing nodes, and the
automatically generated version did not have any noticeable overheads
compared to hand written codes. The speedup obtained by using GPU over
using only the CPU on each node of a cluster is between 3 and 21.

#e1#

#s1#
#abstract1250.txt#


This paper presents a support function for MPI derived datatypes on an
enhancer of memory and network named DIMMnet-3. It is a network
interface with vector access functions and multi-banked extended
memory, which is under development. Semi-hardwired derived datatype
communication based on RDMA with hardwired scatter and gather is
proposed. This mechanism and MPI using it are implemented and validated
on DIMMnet-2 which is a former prototype operating on DDR DIMM slot.
The performance of scatter and gather transfer of 8byte elements with
large interval by using vector commands of DIMMnet-2 is 6.8 compared
with software on a host. Proprietary benchmark of MPI derived datatype
communication for transferring a submatrix corresponding to a narrow
HALO area is executed. Observed bandwidth on DIMMnet-2 is far higher
than that for similar condition with VAPI based MPI implementation on
InfniBand, even though very old generation FPGA, poorer CPU and
motherboard are used. This function will avoid cache pollution and save
CPU time for processing with local data which can be overlapped with
communication. A new commercial machine with vector scatter/gather
functions in NIC named SGI Altix UV is launched recently. It may be
able to adopt our proposed concept partially, even though the capacity
and fine grain access throughput of main memory attached with CPU are
not enhanced on it.

#e1#

#s1#
#abstract1251.txt#


We propose an algorithm that can improve the quality of the
reconstructed image from the single hologram recorded by the optical
system of the parallel four-step phase-shifting digital holography. The
proposed algorithm applies the image-reconstruction algorithm of
parallel two-step phase-shifting digital holography to the hologram so
as to reduce errors in the reconstructed image and eliminate ghosts. We
numerically and experimentally confirmed that the proposed algorithm
decreased 25% in terms of root mean square error in amplitude, and
eliminated the ghosts, respectively.

#e1#

#s1#
#abstract1252.txt#


For MIMO systems, the ST-BICM approach using iterative processing has
been recognized as a method for achieving near-capacity performance.
However, the a posteriori probability calculator in the MIMO detector,
relying on exhaustive or partial search of candidate bit vectors, is
not amenable to practical implementation at high rates ( #or= 16 raw
bits per channel use) due its exponential complexity in rate. On the
other hand, recently developed low complexity single stream demappers
based on a parallel approach, while yielding comparable performance in
ideal conditions, suffer significant performance loss in several
practical scenarios. In this paper, we propose a novel demapper which
closes the performance gap between the low complexity detectors based
on single stream demapping and their exponentially complex counterparts.

#e1#

#s1#
#abstract1253.txt#


In this paper, the high-efficient and reconfigurable architectures for
the 9/7-5/3 discrete wavelet transform (DWT) based on convolution
scheme are proposed. The proposed parallel and pipelined architectures
consist of a high-pass filter (HF) and a low-pass filter (LF). The
critical paths of the proposed architectures are reduced. Filter
coefficients of the biorthogonal 9/7-5/3 wavelet low-pass filter are
quantized before implementation in the high-speed computation hardware.
In the proposed architectures, all multiplications are performed using
less shifts and additions. The proposed reconfigurable architecture is
100% hardware utilization and ultra low-power. The proposed
reconfigurable architectures have regular structure, simple control
flow, high throughput and high scalability. Thus, they are very
suitable for new-generation image compression systems, such as
IPEG-2000.

#e1#

#s1#
#abstract1254.txt#


In this paper, the high-efficient and reconfigurable lined-based
architectures for the 9/7-5/3 discrete wavelet transform (DWT) based on
lifting scheme are proposed. The proposed parallel and pipelined
architectures consist of a horizontal filter (HF) and a vertical filter
(VF). The critical paths of the proposed architectures are reduced.
Filter coefficients of the biorthogonal 9/7-5/3 wavelet low-pass filter
are quantized before implementation in the high-speed computation
hardware In the proposed architectures, all multiplications are
performed using less shifts and additions. The proposed reconfigurable
architecture is 100% hardware utilization and ultra low-power. The
proposed reconfigurable architectures have regular structure, simple
control flow, high throughput and high scalability. Thus, they are very
suitable for new-generation image compression systems, such as
JPEG-2000.

#e1#

#s1#
#abstract1255.txt#


The automatic detection and classification of manmade objects in
overhead imagery is key to generating geospatial intelligence (GEOINT)
from today's high space-time bandwidth sensors in a timely manner. A
flexible multi-stage object detection and classification capability
known as the IMINT Data Conditioner (IDC) has been developed that can
exploit different kinds of imagery using a mission-specific processing
chain. A front-end data reader/tiler converts standard imagery products
into a set of tiles for processing, which facilitates parallel
processing on multiprocessor/multithreaded systems. The first stage of
processing contains a suite of object detectors designed to exploit
different sensor modalities that locate and chip out candidate object
regions. The second processing stage segments object regions, estimates
their length, width, and pose, and determines their geographic
location. The third stage classifies detections into one of K
predetermined object classes (specified in a models file) plus clutter.
Detections are scored based on their salience, size/shape, and
spatial-spectral properties. Detection reports can be output in a
number of popular formats including flat files, HTML web pages, and KML
files for display in Google Maps or Google Earth. Several examples
illustrating the operation and performance of the IDC on Quickbird,
GeoEye, and DCS SAR imagery are presented.

#e1#

#s1#
#abstract1256.txt#


Recently, GPU computing has taken the scientific computing landscape by
storm, fueled by the attractive nature of the massively parallel
arithmetic hardware. When porting their code, researchers rely on a set
of best practices that have been developed over the few years that
general purpose GPU computing has been employed. This paper challenges
a widely held belief that transfers to and from the GPU device must be
minimized to achieve the best speedups over existing codes by
presenting a case study on CULA, our library for dense linear algebra
computation on GPU. Among the topics to be discussed include the
relationship between computation and transfer time for both synchronous
and asynchronous transfers, as well as the impact that data allocations
have on memory performance and overall solution time.

#e1#

#s1#
#abstract1257.txt#


The modern graphics processing unit (GPU) found in many standard
personal computers is a highly parallel math processor capable of
nearly 1 TFLOPS peak throughput at a cost similar to a high-end CPU and
an excellent FLOPS/watt ratio. High-level linear algebra operations are
computationally intense, often requiring O(N3) operations and would
seem a natural fit for the processing power of the GPU. Our work is on
CULA, a GPU accelerated implementation of linear algebra routines. We
present results from factorizations such as LU decomposition, singular
value decomposition and QR decomposition along with applications like
system solution and least squares. The GPU execution model featured by
NVIDIA GPUs based on CUDA demands very strong parallelism, requiring
between hundreds and thousands of simultaneous operations to achieve
high performance. Some constructs from linear algebra map extremely
well to the GPU and others map poorly. CPUs, on the other hand, do well
at smaller order parallelism and perform acceptably during
low-parallelism code segments. Our work addresses this via hybrid a
processing model, in which the CPU and GPU work simultaneously to
produce results. In many cases, this is accomplished by allowing each
platform to do the work it performs most naturally.

#e1#

#s1#
#abstract1258.txt#


We have proposed a data acquisition system with high speed USB
interface using FPGA chip as the main processing unit. Since the FPGA
has a number of modules on chip, which can operate independently, it
can be utilized for the data acquisition system with multi-channels for
the connection to four ADC signals with four different protocols of
Parallel, SPI, I/sup 2/C and one-wire protocol. The system is
controlled by the software written in the visual C/sup ++/. It allows
the user to be able to interface to a PC for data restoration and
monitoring. We found that this system can perform data acquisition with
high rate data transfer.

#e1#

#s1#
#abstract1259.txt#


A conventional pseudorandom sequence generator creates only 1 bit of
data per clock cycle. Therefore, it may cause a delay in data
communications. In this paper, we propose an efficient implementation
method for a pseudorandom sequence generator with parallel outputs. By
virtue of the simple matrix multiplications, we derive a well-organized
recursive formula and realize a pseudorandom sequence generator with
multiple outputs. Experimental results show that, although the total
area of the proposed scheme is 3% to 13% larger than that of the
existing scheme, our parallel architecture improves the throughput by
2,4, and 6 times compared with the existing scheme based on a single
output In addition, we apply our approach to a 2*2 multiple
input/multiple output (MEMO) detector targeting the 3rd Generation
Partnership Project Long Term Evolution (3GPP LTE) system. Therefore,
the throughput of the MEMO detector is significantly enhanced by
parallel processing of data communications.

#e1#

#s1#
#abstract1260.txt#


It is important to agricultural production to estimate a growth state
of rice in paddy field. The system that diagnoses the growth of rice by
the image analysis is proposed. This system transforms bird's-eye view
images of paddy fields to top views of a parallel projection and
analyzes them. Image analysis is performed for each image. Therefore,
there is a possibility that we cannot diagnose all farms within a
limited period in case a farm to diagnose is wide. Analytical time is
shortened by composing plural plane views to one piece of image. In
this paper, transformation and composition methods of field pictures
are proposed. The angle of depression of the camera is automatically
calculated from the field picture and the map, and the field images are
transformed to a parallel projected image. Parallel projected field
image of target area is made by transforming them to be suitable for
the map, and composing it.

#e1#

#s1#
#abstract1261.txt#


After the mechanism of device driver based on RTOS of VxWorks is
analyzed, the device driver of 82C55A parallel interface is implemented
by connect the application layer and device driver through I/O system
technique. Due to the codes of VxWorks system is not open, combining
with the need of project and the core codes including read, write,
choice control of the driver program are innovative offered and the
method to design software of application layer is illuminated, the
detailed artifice of measurement under the pc104 hardware is put out
also, the test illuminates that the device driver is right and method
of test is applicable, it has good value on project.

#e1#

#s1#
#abstract1262.txt#


Problem statement: Calculating sensitive functions for a large
dimension control system to find the unknowns vectors for a linear
system in both single and multi processors, is not considered
internally compatible with multi tasking environments, so breaking the
process can cost time and memory and it couldn't be paused, resumed and
saved as patterns for later continuity. This study is an attempt to
solve this problem in parallel to reduce the time factor needed and
increase the efficiency by using parallel calculation sensitivity
function for multi tasking environments (PSME) algorithm. Approach:
calculate in parallel sensitivity function using n-1 processors where n
is a number of linear equations which can be represented as TX = W,
where T is a matrix of size n/sub 1/*n/sub 2/, X = T/sup -1/W, is a
vector of unknowns and part X/ part h = T/sup -1/(( part T/ part h)-(
part W/ part h)) is a sensitivity function with respect to variation of
system components h. The algorithm (PSME) divides the mathematical
input model into two partitions and uses only (n-1) processors to find
the vector of unknowns for original system x = (x/sub 1/,x/sub
2/,...,x/sub n/)/sup T/ and in parallel using (n-1) processors to find
the vector of unknowns for similar system (x')/sup t/ = d/sup t/T/sup
-1/ = (x/sub 1/',X/sub 2/',...x/sub n/') tau by using Net-Processors,
where d is a constant vector. Finally, sensitivity function with
respect to variation of component part X/ part h/sub i/ = (x/sub
X/*x/sub i/)can be calculated in parallel by multiplication unknowns
X/sub i/*X/sub i/, where i = 0,1,...n-1. Results: The running time t is
reduced to 0(t/n-l) and, the performance of (PSME) was increased by
30-40%. Conclusion: Hence, used (PSME) algorithm reduced the time to
calculate sensitivity function for a large dimension control system and
the performance was increased.

#e1#

#s1#
#abstract1264.txt#


Many applications need efficient eye feature extraction algorithms.
This paper presents a modified parabolic Hough transform where the
coefficient of the parabola is used as the variant with the vertex
locations changing as the coefficient changes. The gradient direction
of edge feature is also combined to reduce the computational cost. The
algorithm is also extended to a parabola in which the axis of symmetry
is not parallel to the coordinate axes. Experimental results show that
the method can extract eye feature precisely and efficiently, and can
be applied in the areas such as human-computer interface and driver
fatigue detection.

#e1#

#s1#
#abstract1265.txt#


In this paper, we present a fast Fourier transform (FFT) processor with
four parallel data paths for multiband orthogonal frequency-division
multiplexing ultrawideband systems. The proposed 128-point FFT
processor employs both a modified radix-24 algorithm and a radix-23
algorithm to significantly reduce the numbers of complex constant
multipliers and complex booth multipliers. It also employs
substructure-sharing multiplication units instead of constant
multipliers to efficiently conduct multiplication operations with only
addition and shift operations. The proposed FFT processor is
implemented and tested using 0.18 mu m CMOS technology with a supply
voltage of 1.8 V. The hardware-efficient 128-point FFT processor with
four data streams can support a data processing rate of up to 1
Gsample/s while consuming 112 mW. The implementation results show that
the proposed 128-point mixed-radix FFT architecture significantly
reduces the hardware cost and power consumption in comparison to
existing 128-point FFT architectures.

#e1#

#s1#
#abstract1266.txt#


This paper proposes a multi-query optimization algorithm for
pipeline-based distributed similarity query processing (pGMSQ) in grid
environment. First, when a number of query requests are simultaneously
submitted by users, a cost-based dynamic query clustering (DQC) is
invoked to quickly and effectively identify the correlation among the
query spheres (requests). Then, index-support vector set reduction is
performed at data node level in parallel. Finally, refinement of the
candidate vectors is conducted to get the answer set at the execution
node level. By adopting pipeline-based technique, this algorithm is
experimentally proved to be efficient and effective in minimizing the
response time by decreasing network transfer cost and increasing the
throughput.

#e1#

#s1#
#abstract1273.txt#


We study the interaction between the MIMD (Multiplicative Increase
Multiplicative Decrease) congestion control and a bottleneck router
with Drop Tail buffer. We consider the problem in the framework of
deterministic hybrid models. We study conditions under which the system
trajectories converge to limiting cycles with a single jump. Following
that, we consider the problem of the optimal buffer sizing in the
framework of multi-criteria optimization in which the Lagrange function
corresponds to a linear combination of the average throughput and the
average delay in the queue. As case studies, we consider the Slow Start
phase of TCP New Reno and Scalable TCP for high speed networks. [All
rights reserved Elsevier].

#e1#

#s1#
#abstract1293.txt#


The aim of the paper is to validate a software architecture that allows
an image processing researcher to develop parallel applications. The
challenge was to develop algorithms that perform real-time low level
operations on digital images able to be executed on a cluster of
desktop PCs. The experiments show how to use parallelizable patterns
and how to optimize the load balancing between the workstations.

#e1#

#s1#
#abstract1294.txt#


In this paper we discuss a new model for document clustering which has
been adapted using non-negative matrix factorization method. The key
idea is to cluster the documents after measuring the proximity of the
documents with the extracted features. The extracted features are
considered as the final cluster labels and clustering is done using
cosine similarity which is equivalent to k-means with a single turn.
This model was implemented using apache lucene project for indexing
documents and mapreduce framework of apache hadoop project for parallel
implementation of k-means algorithm. Since experiments were carried
only in one cluster of Hadoop, the significant reduction in time was
obtained by mapreduce implementation when clusters size exceeded 9 i.e.
40 documents averaging 1.5 kilobytes. Thus it is concluded that the
feature extracted using NMF can be used to cluster documents
considering them to be final cluster labels as in k-means, and for
large scale documents, the parallel implementation using mapreduce can
lead to reduction of computational time. We have termed this model as
KNMF (K-means with NMF algorithm).

#e1#

#s1#
#abstract1298.txt#


Fully exploiting the spatial feature of image makes H.264/ AVC standard
superior in intra prediction part. However, when hardware is
considered, full support of all intra modes will cause high design
effort, especially for large image size. In this paper, we propose a
low design effort solution for intra predictor generation, which is the
most significant part in intra engine. Firstly, one parallel processing
flow is given out, which achieves 37.5% reduction of processing time.
Secondly, a fully utilized predictor generation architecture is given
out, which saves 77.5% cycles of original one. With 30.11k gates at
200MHz, our design can support full-mode intra prediction for real-time
processing of 4k * 2k@60fps.

#e1#

#s1#
#abstract1307.txt#


The authors present a novel approach of using reconfigurable fabric to
accelerate a face detection algorithm based on the Haar classifier.
With highly pipelined architecture and utilising abundant parallel
arithmetic units in FPGA, the authors have achieved real-time
performance of face detection with very high detection rate and low
false positives. The 1-classifier and 16-classifier realisations in an
accelerator provide 10* and 72* speedups, respectively, over the
software counterpart. Moreover, the authors*, approach is scalable
towards the resources available on FPGA and it will gain more momentum
as the Geneseo Initiative is introduced in the market. This work also
provides an understanding of using the reconfigurable fabric for
accelerating non-systolic-based vision algorithms.

#e1#

#s1#
#abstract1308.txt#


An efficient micromagnetic solver running on graphics processing units
(GPU) is demonstrated. The solver implements a nonuniform grid
interpolation method (NGIM) to compute the superposition integral for
the magnetostatic field with operations and memory requirements. The
NGIM divides the computational domain into a hierarchy of boxes
containing sources and observers, and it uses spatial interpolation
from sparse nonuniform grids to achieve computational savings.
Efficiency of the GPU solver is achieved by using coalesced memory
accessing requiring arranging data in contiguous addresses,
one-block-per-box computations with a block of threads handling an
observation box to achieve the best utilization of the GPU threads, and
on-fly computation of all grids and interpolation coefficients leading
to reduced memory and increased speed. The GPU-CPU speed-ups are shown
to be in the range 40-100 depending on the problem size and accuracy. A
simple and inexpensive GPU is shown to handle efficiently problems
comprising discretizations of more than 16 million of spins.

#e1#

#s1#
#abstract1309.txt#


We have adapted our finite element micromagnetic simulation software to
the massively parallel architecture of graphical processing units
(GPUs) with double-precision floating point accuracy. Using the example
of Standard Problem #4 with different numbers of discretization points,
we demonstrate the high speed performance of a single GPU compared with
an OpenMP-parallelized version of the code using eight CPUs. The
adaption of both the magnetostatic field calculation and the time
integration of the Landau-Lifshitz-Gilbert equation routines can lead
to a speedup factor of up to four. The gain in computation performance
of the GPU code increases with increasing number of discretization
nodes. The computation time required for high-resolution micromagnetic
simulations of the magnetization dynamics in large magnetic samples can
thus be reduced effectively by employing GPUs.

#e1#

#s1#
#abstract1326.txt#


Background: MapReduce is a parallel framework that has been used
effectively to design largescale parallel applications for large
computing clusters. In this paper, we evaluate the viability of the
MapReduce framework for designing phylogenetic applications. The
problem of interest is generating the all-to-all Robinson-Foulds
distance matrix, which has many applications for visualizing and
clustering large collections of evolutionary trees. We introduce MrsRF
(MapReduce Speeds up RF), a multi-core algorithm to generate a t * t
Robinson-Foulds distance matrix between t trees using the MapReduce
paradigm. Results: We studied the performance of our MrsRF algorithm on
two large biological trees sets consisting of 20,000 trees of 150 taxa
each and 33,306 trees of 567 taxa each. Our experiments show that MrsRF
is a scalable approach reaching a speedup of over 18 on 32 total cores.
Our results also show that achieving top speedup on a multi-core
cluster requires different cluster configurations. Finally, we show how
to use an RF matrix to summarize collections of phylogenetic trees
visually. Conclusion: Our results show that MapReduce is a promising
paradigm for developing multicore phylogenetic applications. The
results also demonstrate that different multi-core configurations must
be tested in order to obtain optimum performance. We conclude that RF
matrices play a critical role in developing techniques to summarize
large collections of trees.

#e1#

#s1#
#abstract1327.txt#


A novel technique for multiple-image optical encryption is proposed, in
which a set of parallel plaintexts can be extracted from the same
designed ciphertext respectively. In the process of encryption, the
principle of random phase encoding is utilized, and the phase keys
corresponding to different plaintexts are achieved independently from
the same designed ciphertext by cascade phase retrieval algorithm
(CPRA). The advantages of the approach could be concluded as
implementing decryption without cross-talk, infinite encrypted capacity
and simple architecture. And the plaintexts extracted mode is extended
from peer-to-peer to peer-to-multipeer. Numerical simulation verifies
the validity. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1328.txt#


We present a number of optimization techniques to compute prefix sums
on linked lists and implement them on multithreaded GPUs using CUDA.
Prefix computations on linked structures involve in general highly
irregular fine grain memory accesses that are typical of many
computations on linked lists, trees, and graphs. While the current
generation of GPUs provides substantial computational power and
extremely high bandwidth memory accesses, they may appear at first to
be primarily geared toward streamed, highly data parallel computations.
In this paper, we introduce an optimized multithreaded GPU algorithm
for prefix computations through a randomization process that reduces
the problem to a large number of fine-grain computations. We map these
fine-grain computations onto multithreaded GPUs in such a way that the
processing cost per element is shown to be close to the best possible.
Our experimental results show scalability for list sizes ranging from
1M nodes to 256M nodes, and significantly improve on the recently
published parallel implementations of list ranking, including
implementations on the Cell Processor, the MTA-8, and the NVIDIA
GeForce 200 series. They also compare favorably to the performance of
the best known CUDA algorithm for the scan operation on the Tesla C1060.

#e1#

#s1#
#abstract1329.txt#


To exploit the potential of multicore architectures, recent dense
linear algebra libraries have used tile algorithms, which consist in
scheduling a Directed Acyclic Graph (DAG) of tasks of fine granularity
where nodes represent tasks, either panel factorization or update of a
block-column, and edges represent dependencies among them. Although
past approaches already achieve high performance on moderate and large
square matrices, their way of processing a panel in sequence leads to
limited performance when factorizing tall and skinny matrices or small
square matrices. We present a new fully asynchronous method for
computing a QR factorization on shared-memory multicore architectures
that overcomes this bottleneck. Our contribution is to adapt an
existing algorithm that performs a panel factorization in parallel
(named Communication-A voiding QR and initially designed for
distributed-memory machines), to the context of tile algorithms using
asynchronous computations. An experimental study shows significant
improvement (up to almost 10 times faster) compared to state-of-the-art
approaches. We aim to eventually incorporate this work into the
Parallel Linear Algebra for Scalable Multi-core Architectures (PLASMA)
library.

#e1#

#s1#
#abstract1330.txt#


Adaptive resource allocation with different numbers of machine nodes
provides more flexibility and significantly better potential
performance for local job and grid scheduling. With the emergence of
parallel computing in every-day life on multi-core systems, such
schedulers will likely increase in practical relevance. A major reason
why adaptive schedulers are not yet practically used is lacking
knowledge of the scalability curves of the applications. Existing
white-box approaches for scalability prediction are too expensive to
apply them routinely. We present ADEPT, a speedup and runtime
prediction tool, which is inexpensive and easy-to-use. ADEPT employs a
black-box model and can be practically applied at large scale without
user or administrator involvement. ADEPT requires neither program
analysis and measurements nor user guesses but makes highly accurate
predictions with only few observations of application runtime over
different numbers of nodes/cores. ADEPT performs efficient model
fitting by introducing an envelope-derivation technique to constrain
the search. Additionally, ADEPT is capable of handling deviations from
the underlying model by detection and automatic correction of anomalies
via a fluctuation metric and by considering specific scalability
patterns via multi-phase modeling. ADEPT also performs reliability
judgment with potential proposal for placement of additional
observations. Using MPI and OpenMP implementations of the NAS
benchmarks and seven real applications, we demonstrate the
effectiveness and high prediction accuracy of ADEPT for both speedup
and runtime prediction, including interpolative and extrapolative
cases, and show the capability of ADEPT to successfully handle special
cases.

#e1#

#s1#
#abstract1331.txt#


Document clustering plays an important role in data mining systems.
Recently, a flocking-based document clustering algorithm has been
proposed to solve the problem through simulation resembling the
flocking behavior of birds in nature. This method is superior to other
clustering algorithms, including k-means, in the sense that the outcome
is not sensitive to the initial state. One limitation of this approach
is that the algorithmic complexity is inherently quadratic in the
number of documents. As a result, execution time becomes a bottleneck
with large number of documents. In this paper, we assess the benefits
of exploiting the computational power of Beowulf-like clusters equipped
with contemporary Graphics Processing Units (GPUs) as a means to
significantly reduce the runtime of flocking-based document clustering.
Our framework scales up to over one million documents processed
simultaneously in a sixteen-node moderate GPU cluster. Results are also
compared to a four-node cluster with higher-end GPUs. On these
clusters, we observe 30X-50X speedups, which demonstrate the potential
of GPU clusters to efficiently solve massive data mining problems. Such
speedups combined with the scalability potential and accelerator-based
parallelization are unique in the domain of document-based data mining,
to the best of our knowledge.

#e1#

#s1#
#abstract1332.txt#


Sorting is a commonly used process with a wide breadth of applications
in the high performance computing field. Early research in parallel
processing has provided us with comprehensive analysis and theory for
parallel sorting algorithms. However, modern supercomputers have
advanced rapidly in size and changed significantly in architecture,
forcing new adaptations to these algorithms. To fully utilize the
potential of highly parallel machines, tens of thousands of processors
are used. Efficiently scaling parallel sorting on machines of this
magnitude is inhibited by the communication-intensive problem of
migrating large amounts of data between processors. The challenge is to
design a highly scalable sorting algorithm that uses minimal
communication, maximizes overlap between computation and communication,
and uses memory efficiently. This paper presents a scalable extension
of the Histogram Sorting method, making fundamental modifications to
the original algorithm in order to minimize message contention and
exploit overlap. We implement Histogram Sort, Sample Sort, and Radix
Sort in Charm++ and compare their performance. The choice of algorithm
as well as the importance of the optimizations is validated by
performance tests on two predominant modern supercomputer
architectures: XT4 at ORNL (Jaguar) and Blue Gene/P at ANL (Intrepid).

#e1#

#s1#
#abstract1333.txt#


Dictionary-based string matching (DBSM) is a critical component of Deep
Packet Inspection (DPI), where thousands of malicious patterns are
matched against high-bandwidth network traffic. Deterministic finite
automata constructed with the Aho-Corasick algorithm (AC-DFA) have been
widely used for solving this problem. However, the state transition
table (STT) of a large-scale DBSM AC-DFA can span hundreds of megabytes
of system memory, whose limited bandwidth and long latency could become
the performance bottleneck We propose a novel partitioning algorithm
which converts an AC-DFA into a "head" and a "body" parts. The head
part behaves as a traditional AC-DFA that matches the pattern prefixes
up to a predefined length; the body part extends any head match to the
full pattern length in parallel body-tree traversals. Taking advantage
of the SIMD instructions in modern x86-64 multi-core processors, we
design compact and efficient data structures packing multi-path and
multi-stride pattern segments in the body-tree. Compared with an
optimized AC-DFA solution, our head-body matching (HBM) implementation
achieves 1.2x to 3x throughput performance when the input match
(attack) ratio varies from 2% to 32%, respectively. Our HBM data
structure is over 20x smaller than a fully-populated AC-DFA for both
Snort and ClamAV dictionaries. The aggregated throughput of our HBM
approach scales almost 7x with 8 threads to over 10 Gbps in a
dual-socket quad-core Opteron (Shanghai) server.

#e1#

#s1#
#abstract1334.txt#


Ensuring the correctness of complex implementations of software
transactional memory (STM) is a daunting task. Attempts have been made
to formally verify STMs, but these are limited in the scale of systems
they can handle and generally verify only a model of the system, and
not the actual system. In this paper we present an alternate attack on
checking the correctness of an STM implementation by verifying the
execution runs of an STM using a checker that runs in parallel with the
transaction memory system. With future many-core systems predicted to
have hundreds and even thousands of cores, it is reasonable to utilize
some of these cores for ensuring the correctness of the rest of the
system. This will be needed anyway given the increasing likelihood of
dynamic errors due to particle hits (soft errors) and increasing
fragility of nanoscale devices. These errors can only be detected at
runtime. An important correctness criterion that is the subject of
verification is the serializability of transactions. While checking
transaction serializability is NP-complete, practically useful
subclasses such as interchange-serializability (DSR) are efficiently
computable. Checking DSR reduces to checking for cycles in a
transaction ordering graph which captures the access order of objects
shared between transaction instances. Doing this concurrent to the main
transaction execution requires minimizing the overhead of capturing
object accesses, and managing the size of the graph, which can be as
large as the total number of dynamic transactions and object accesses.
We discuss techniques for minimizing the overhead of access logging
which includes time-stamping, and present techniques for on-the-fly
graph compaction that drastically reduce the size of the graph that
needs to be maintained, to be no larger than the number of threads. We
have implemented concurrent serializability checking in the Rochester
Software Transactional Memory (RSTM) system. We present our practical
experiences with this including results for the RSTM, STAMP and
synthetic benchmarks. The overhead of concurrent checking is a strong
function of the transaction length. For long transactions this is
negligible. Thus the use of the proposed method for continuous runtime
checking is acceptable. For very short transactions this can be
significant. In this case we see the applicability of the proposed
method for debugging.

#e1#

#s1#
#abstract1335.txt#


The following topics are dealt with: parallel processing; distributed
processing; network management; scientific computing; GPU; data storage
system; data memory system; fault tolerance; sorting; performance
improvement; scalability improvement; network architecture;
benchmarking tools; resource allocation; image processing; data mining;
transactional memory; correctness analysis; parallel linear algebra;
P2P algorithm; string problem; sequence problem; energy-aware task
management; parallel operating system; parallel graph; caching; thread
scheduling; automatic tuning; automatic parallelization; architectural
support; runtime system; client-server system management; wireless
network; and data management.

#e1#

#s1#
#abstract1336.txt#


Heterogeneity complicates the efficient use of multicomputer platforms,
but does it enhance their performance? their cost effectiveness? How
can one measure the power of a heterogeneous assemblage of computers
("cluster," for short), both in absolute terms (how powerful is this
cluster) and relative terms (which cluster is more powerful)? What
makes one cluster more powerful than another? Is one better off with a
cluster that has one super-fast computer and the rest of "average"
speed or with a cluster all of whose computers are "moderately" fast?
If you could replace just one computer in your cluster with a faster
one, which computer would you choose: the fastest? the slowest? How
does one even ask questions such as these in a rigorous, yet tractable
manner? A framework is proposed, and some answers are derived, a few
rather surprising. Three highlights: (1) If one can replace only one
computer in a cluster by a faster one, it is provably (almost) always
most advantageous to replace the fastest one. (2) If the computers in
two clusters have the same mean speed, then, empirically, the cluster
with the larger variance in speed is (almost) always the faster one.
(3) Heterogeneity can actually lend power to a cluster!.

#e1#

#s1#
#abstract1337.txt#


Line-rate data traffic monitoring in high-speed networks is essential
for network management. To satisfy the line-rate requirement, one can
leverage multi-core architectures to parallelize traffic monitoring so
as to improve information processing capabilities over traditional
uni-processor architectures. Nevertheless, realizing the full potential
of multi-core architectures still needs substantial work, especially in
the face of the ever-increasing volume and complexity of network
traffic. This paper addresses the issue through the design of a
lock-free, cache-efficient synchronization mechanism that serves as a
basic building block for a general class of multithreaded, multi-core
traffic monitoring applications. We embed the synchronization mechanism
into MCRingBuffer, a multi-core shared ring buffer that provides fast
data accesses among threads running in different cores. MCRingBuffer
allows concurrent lock-free data accesses and improves the cache
locality of accessing the control variables that are used for thread
synchronization. Through extensive evaluation on an Intel Xeon
multi-core machine, we show that MCRingBuffer achieves a throughput
gain of up to 5* over existing lock-free ring buffers. Finally, we
present a parallel traffic monitoring prototype that is built upon
MCRingBuffer, and demonstrate via trace-driven simulation how
MCRingBuffer facilitates packet processing at line rate.

#e1#

#s1#
#abstract1338.txt#


Innovative scientific applications and emerging dense data sources are
creating a data deluge for high-end computing systems. Processing such
large input data typically involves copying (or staging) onto the
supercomputer's specialized high-speed storage, scratch space, for
sustained high I/O throughput. The current practice of conservatively
staging data as early as possible makes the data vulnerable to storage
failures, which may entail re-staging and consequently reduced job
throughput. To address this, we present a timely staging framework that
uses a combination of job startup time predictions, user-specified
intermediate nodes, and decentralized data delivery to coincide input
data staging with job start-up. By delaying staging to when it is
necessary, the exposure to failures and its effects can be reduced.
Evaluation using both PlanetLab and simulations based on three years of
Jaguar (No. 1 in Top500) job logs show as much as 85.9% reduction in
staging times compared to direct transfers, 75.2% reduction in wait
time on scratch, and 2.4% reduction in usage/hour.

#e1#

#s1#
#abstract1339.txt#


With larger and larger systems being constantly deployed, trace-based
performance analysis of parallel applications has become a challenging
task. Even if the amount of performance data gathered per single
process is small, traces rapidly become unmanageable when merging
together the information collected from all processes. In general, an
efficient analysis of such a large volume of data is subject to a
previous filtering step that directs the analyst's attention towards
what is meaningful to understand the observed application behavior.
Furthermore, the iterative nature of most scientific applications
usually ends up producing repetitive information. Discarding irrelevant
data aims at reducing both the size of traces, and the time required to
perform the analysis and deliver results. In this paper, we present an
on-line analysis framework that relies on clustering techniques to
intelligently select the most relevant information to understand how
the application behaves, while keeping the volume of performance data
at a reasonable size.

#e1#

#s1#
#abstract1340.txt#


We consider the problem of scheduling jobs on a pool of machines. Each
job requires multiple machines on which it executes in parallel. For
each job, the input specifies release time, deadline, processing time,
profit and the number of machines required. The total number of
machines may be different at different points of time. A feasible
solution is a subset of jobs and a schedule for them such that at any
timeslot, the total number of machines required by the jobs active at
the timeslot does not exceed the number of machines available at that
timeslot. We present an O(log(B/sub max//B/sub min/))-approximation
algorithm, where B/sub max/ and B/sub min/ are the maximum and minimum
available bandwidth (maximum and minimum number of machines available
over all the timeslots). Our algorithm and the approximation ratio are
applicable for more a general problem that we call the Varying
bandwidth resource allocation problem with bag constraints (BAGVBRAP).
The BAGVBRAP problem is a generalization of some previously studied
scheduling and resource allocation problems.

#e1#

#s1#
#abstract1341.txt#


Pipelined workflows are a popular programming paradigm for parallel
applications. In these workflows, the computation is divided into
several stages, and these stages are connected to each other through
first-in first-out channels. In order to execute these workflows on a
parallel machine, we must first determine the mapping of the stages
onto the various processors on the machine. After finding the mapping,
we must compute the schedule, i.e., the order in which the various
stages execute on their assigned processors. In this paper, we assume
that the mapping is given and explore the latter problem of scheduling,
particularly for linear workflows. Linear workflows are those in which
dependencies between stages can be represented by a linear graph. The
objective of the scheduling algorithm is either to minimize the period
(the inverse of the throughput), or to minimize the latency (response
time), or both. We consider two realistic execution models: the
one-port model (all operations are serialized) and the multi-port model
(bounded communication capacities and communication/computation
overlap). In both models, finding a schedule to minimize the latency is
easy. However, computing the schedule to minimize the period is NP-hard
in the one-port model, but can be done in polynomial time in the
multi-port model. We also present an approximation algorithm to
minimize the period in the one-port model. Finally, the bi-criteria
problem, which consists in finding a schedule respecting a given period
and a given latency, is NP-hard in both models.

#e1#

#s1#
#abstract1342.txt#


Determinant Quantum Monte Carlo (DQMC) simulation has been widely used
to reveal macroscopic properties of strong correlated materials.
However, parallelization of the DQMC simulation is extremely
challenging duo to the serial nature of underlying Markov chain and
numerical stability issues. We extend previous work with novelty by
presenting a hybrid granularity parallelization (HGP) scheme that
combines algorithmic and implementation techniques to speed up the DQMC
simulation. From coarse-grained parallel Markov chain and task
decompositions to fine-grained parallelization methods for matrix
computations and Green's function calculations, the HGP scheme explores
the parallelism on different levels and maps the underlying algorithms
onto different computational components that are suitable for modern
high performance heterogeneous computer systems. Practical techniques,
such as communication and computation overlapping, message compression
and load balancing are also considered in the proposed HGP scheme. We
have implemented the DQMC simulation with the HGP scheme on an IBM Blue
Gene/P system. The effectiveness of the new scheme is demonstrated
through both theoretical analysis and performance results. Experiments
have shown over a factor of 80 speedups on an IBM Blue Gene/P system
with 1,014 computational processors.

#e1#

#s1#
#abstract1343.txt#


In this paper, we study the problem of finding optimal mappings for
several independent but concurrent workflow applications, in order to
optimize performance-related criteria together with energy consumption.
Each application consists in a linear chain graph with several stages,
and processes successive data sets in pipeline mode, from the first to
the last stage. We study the problem complexity on different target
execution platforms, ranking from fully homogeneous platforms to fully
heterogeneous ones. The goal is to select an execution speed for each
processor, and then to assign stages to processors, with the aim of
optimizing several concurrent optimization criteria. There is a clear
trade-off to reach, since running faster and/or more processors leads
to better performance, but the energy consumption is then very high.
Energy savings can be achieved at the price of a lower performance, by
reducing processor speeds or enrolling fewer resources. We consider two
mapping strategies: in one-to-one mappings, a processor is assigned a
single stage, while in interval mappings, a processor may process an
interval of consecutive stages of the same application. For both
mapping strategies and all platform types, we establish the complexity
of several multi-criteria optimization problems, whose objective
functions combine period, latency and energy criteria. In particular,
we exhibit cases where the problem is NP-hard with concurrent
applications, while it can be solved in polynomial time for a single
application. Also, we demonstrate the difficulty of performance/energy
trade-offs by proving that the tri-criteria problem is NP-hard, even
with a single application on a fully homogeneous platform.

#e1#

#s1#
#abstract1344.txt#


We present an implementation of one of the direct self-consistent-field
(DSCF) calculation techniques, the restricted Hartree-Fock method, on a
high-performance computing cluster outfitted with graphics processing
units (GPUs) and demonstrate its effectiveness and scalability up to
128 cluster nodes on molecules of as many as 1,732 atoms. We discuss
the overall parallel application architecture that relies on message
passing interface for distributing workload among GPU cluster nodes and
POSIX threads to manage the use of GPUs internal to each node. This
approach of combining coarse and fine-grain parallelism on a
distributed memory system allows to perform DSCF calculations on
molecules that up until now have been unattainable due to the excessive
computational requirements.

#e1#

#s1#
#abstract1345.txt#


It is widely believed that most Recognition and Mining (RM) workloads
can easily take advantage of parallel computing platforms because these
workloads are dataparallel. Contrary to this popular belief, we present
RM workloads for which conventional parallel implementations scale
poorly on multi-core platforms. We identify off-chip memory transfers
and overheads in the parallel runtime library as the primary
bottlenecks that limit speedups to be well below the ideal linear
speedup expected for data-parallel workloads. To achieve improved
parallel scalability, we identify and exploit several interesting
properties of RM workloads - sparsity of model updates, low spatial
locality among model updates, presence of insignificant computations,
and the inherently self-healing nature of these algorithms in the
presence of errors. We leverage these domain-specific characteristics
to improve parallel scalability in two major ways. First, we utilize
data dependency relaxation to simultaneously execute multiple training
iterations in parallel, thereby increasing the granularity of the
parallel tasks and significantly lowering the run-time overheads of
fine-grained threading. Second, we strategically drop selected
computations that are insignificant to the accuracy of the final
result, but account for a disproportionately large amount of off-chip
(memory and coherence) traffic. Through the application of the proposed
techniques, we show that much higher speedups are possible on
multi-core platforms for two important RM applications - document
search using semantic indexing, and eye detection in images using
generalized learning vector quantization. On an 8-core platform, we
achieve application speedups of 5.5X and 7.3X compared to sequential
implementations. Compared to conventional parallel implementations of
these applications using Intel's TBB, the proposed techniques result in
4.3X and 4.9X improvements. Although the optimized parallel
implementations are not numerically equivalent to the sequential
implementations, the output quality is shown to be comparable (and
within the margin of variation produced by processing the input data in
a different order). We also explore error mitigation techniques that
can be used to ensure that the accuracy of results is not compromised.

#e1#

#s1#
#abstract1346.txt#


Software Transactional Memory (STM) is a programming paradigm that
allows a programmer to write parallel programs, without having to deal
with the intricacies of synchronization. That burden is instead borne
by the underlying STM system. SwissTM is a lock-based STM, developed at
EPFL, Switzerland. Memory locations map to entries in a lock table to
detect conflicts. Increasing the number of locations that map to a lock
reduces the number of locks to be acquired and improves throughput,
while also increasing the possibility of false conflicts. False
conflicts occur when a transaction that updates a location mapping to a
lock, causes validation failure of another transaction, that reads a
different location mapping to the same lock. In this paper, we present
a solution for the false conflict problem and suggest an adaptive
version of the same algorithm, to improve performance. Our algorithms
produce significant throughput improvement in benchmarks with false
conflicts.

#e1#

#s1#
#abstract1347.txt#


Solving dense linear systems of equations is a fundamental problem in
scientific computing. Numerical simulations involving complex systems
represented in terms of unknown variables and relations between them
often lead to linear systems of equations that must be solved as fast
as possible. We describe current efforts toward the development of
these critical solvers in the area of dense linear algebra (DLA) for
multicore with GPU accelerators. We describe how to code/develop
solvers to effectively use the high computing power available in these
new and emerging hybrid architectures. The approach taken is based on
hybridization techniques in the context of Cholesky, LU, and QR
factorizations. We use a high-level parallel programming model and
leverage existing software infrastructure, e.g. optimized BLAS for CPU
and GPU, and LAPACK for sequential CPU processing. Included also are
architecture and algorithm-specific optimizations for standard solvers
as well as mixed-precision iterative refinement solvers. The new
algorithms, depending on the hardware configuration and routine
parameters, can lead to orders of magnitude acceleration when compared
to the same algorithms on standard multicore architectures that do not
contain GPU accelerators. The newly developed DLA solvers are
integrated and freely available through the MAGMA library.

#e1#

#s1#
#abstract1348.txt#


High-performance computing (HPC) is making its way in every field of
science and engineering by providing advanced methods for getting
deeper comprehension of different processes and phenomena. However, due
to the increased complexity of computer architectures and their
multi-level parallelism, the development of efficient highly parallel
applications is considerably complicated. This process has to be
inevitably augmented with continuous performance analysis in order for
one to be successful in optimizing applications and squeezing the
potential out of today's supercomputers. Periscope is a distributed
performance analysis tool capable of collecting and processing
measurement data from large scale application runs. In comparison to
other similar tools, Periscope provides high-level performance
bottlenecks and not the low-level values of hardware counters. This
paper presents an enhanced and powerful graphical user interface that
was recently developed for Periscope. It was successfully integrated in
the Eclipse development platform as a plug-in and takes advantage of
one of its extensions - the Parallel Tools Platform (PTP). This
approach combines some of platform's advanced programming features with
those of the Periscope performance measurement toolkit. As a result, a
convenient software development and performance analysis environment
was produced that aims at increasing the productivity of developers
during the creation of highly efficient HPC applications.

#e1#

#s1#
#abstract1349.txt#


In this paper, we propose a power-aware parallel job scheduler assuming
DVFS enabled clusters. A CPU frequency assignment algorithm is
integrated into the well established EASY backfilling job scheduling
policy. Running a job at lower frequency results in a reduction in
power dissipation and accordingly in energy consumption. However, lower
frequencies introduce a penalty in performance. Our frequency
assignment algorithm has two adjustable parameters in order to enable
fine grain energy-performance trade-off control. Furthermore, we have
done an analysis of HPC system dimension. This paper investigates
whether having more DVFS enabled processors for same load can lead to
better energy efficiency and performance. Five workload traces from
systems in production use with up to 9 216 processors are simulated to
evaluate the proposed algorithm and the dimensioning problem. Our
approach decreases CPU energy by 7%- 18% on average depending on
allowed job performance penalty. Using the power-aware job scheduling
for 20% larger system, CPU energy needed to execute same load can be
decreased by almost 30% while having same or better job performance.

#e1#

#s1#
#abstract1350.txt#


In this paper, scheduling parallel tasks on multiprocessor computers
with dynamically variable voltage and speed is addressed as
combinatorial optimization problems. Our scheduling problems are
defined such that the energy-delay product is optimized by fixing one
factor and minimizing the other. It is noticed that power-aware
scheduling of parallel tasks has rarely been discussed before. Our
investigation in this paper makes some initial attempt to energy
efficient scheduling of parallel tasks on multiprocessor computers with
dynamic voltage and speed. Our scheduling problems contain three
nontrivial subproblems, namely, system partitioning, task scheduling,
and power supplying. The harmonic system partitioning and processor
allocation scheme is used, which divides a multiprocessor computer into
clusters of equal sizes and schedules tasks of similar sizes together
to increase processor utilization. A three-level energy/time/power
allocation scheme is adopted for a given schedule, such that the
schedule length is minimized by consuming given amount of energy or the
energy consumed is minimized without missing a given deadline. The
performance of our heuristic algorithms is analyzed and accurate
performance bounds are derived. Simulation data which validate our
analytical results are also presented. It is found that our analytical
results provide very accurate estimation of the expected normalized
schedule length and the expected normalized energy consumption, and
that our heuristic algorithms are able to produce solutions very close
to optimum.

#e1#

#s1#
#abstract1351.txt#


Recently, there has been strong interest in large-scale simulations of
biological spiking neural networks (

#e1#

#s1#
#abstract1352.txt#


The increasing availability of multi-core and multiprocessor
architectures provides new opportunities for improving the performance
of many computer simulations. Markov Chain Monte Carlo (MCMC)
simulations are widely used for approximate counting problems, Bayesian
inference and as a means for estimating very high-dimensional
integrals. As such MCMC has found a wide variety of applications in
fields including computational biology and physics, financial
econometrics, machine learning and image processing. Whilst divide and
conquer is an obvious means to simplify image processing tasks,
"naively" dividing an image into smaller images to be processed
separately results in anomalies and breaks the statistical validity of
the MCMC algorithm. We present a method of grouping the spatially local
moves and temporarily partitioning the image to allow those moves to be
processed in parallel, reducing the runtime whilst conserving the
properties of the MCMC method. We calculate the theoretical reduction
in runtime achievable by this method, and test its effectiveness on a
number of different architectures. Experiments are presented that show
reductions in runtime of 38% using a dual-core dual-processor machine.
In circumstances where absolute statistical validity are not required,
an algorithm based upon, but not strictly adhering to, MCMC may
suffice. For such situation two additional algorithms are presented for
partitioning of images to which MCMC will be applied. Assuming
preconditions are met, these methods may be applied with minimal risk
of anomalous results. Although the extent of the runtime reduction will
be data dependent, the experiments performed showed the runtime reduced
to 27% of its original value.

#e1#

#s1#
#abstract1353.txt#


Wild populations of organism are often difficult to study in their
natural settings. Often, it is possible to infer mating information
about these species by genotyping the offspring and using the genetic
information to infer sibling, and other kinship, relationships. While
sibling reconstruction has been studied for a long time, none of the
existing approaches have targeted scalability. In this paper, we
introduce the first parallel approach to reconstructing sibling
relationships from microsatellite markers. We use both functional and
data domain decomposition to break down the problem and argue that this
approach can be applied to other problems where columns are independent
and simple constraint-based enumeration is required. We discuss
algorithmic and implementation choices and their effects on results. We
show that our approach is highly efficient and scalable.

#e1#

#s1#
#abstract1354.txt#


Local search (LS) algorithms are among the most powerful techniques for
solving computationally hard problems in combinatorial optimization.
These algorithms could be viewed as "walks through neighborhoods" where
the walks are performed by iterative procedures that allow to move from
a solution to another one in the solution space. In these heuristics,
designing operators to explore large promising regions of the search
space may improve the quality of the obtained solutions at the expense
of a highly computationally process. Therefore, the use of graphics
processing units (GPUs) provides an efficient complementary way to
speed up the search. However, designing applications on GPU is still
complex and many issues have to be faced. We provide a methodology to
design and implement large neighborhood LS algorithms on GPU. The work
has been experimented for binary problems by deploying multiple
neighborhood structures. The obtained results are convincing both in
terms of efficiency, quality and robustness of the provided solutions
at run time.

#e1#

#s1#
#abstract1355.txt#


Lists intersection is an important operation in modern web search
engines. Many prior studies have focused on the single-core or
multi-core CPU platform or many-core GPU. In this paper, we propose a
CPU-GPU cooperative model that can integrate the computing power of CPU
and GPU to perform lists intersection more efficiently. In the
so-called synchronous mode, queries are grouped into batches and
processed by GPU for high throughput. We design a query-parallel GPU
algorithm based on an element-thread mapping strategy for load
balancing. In the traditional asynchronous model, queries are processed
one-by-one by CPU or GPU to gain perfect response time. We design an
online scheduling algorithm to determine whether CPU or GPU processes
the query faster. Regression analysis on a huge number of experimental
results concludes a regression formula as the scheduling metric. We
perform exhaustive experiments on our new approaches. Experimental
results on the TREC Gov and Baidu datasets show that our approaches can
improve the performance of the lists intersection significantly.

#e1#

#s1#
#abstract1356.txt#


In this work we show how to use a data-dependence profiling tool called
DProf, which can be utilized to assist parallelization for multi-core
systems. DProf is based on an optimizing compiler and uses reference
runs to emit information on runtime dependences between various memory
accesses within a loop. The profiler not only marks the dependent
statements and the accesses but also emits details regarding the
percentage of time the dependences are encountered - the percentage
being taken over the loop iteration size. Though DProf has been
primarily built to capture opportunities for speculative thread-level
parallelism, it has been found that the report generated by the DProf
can be utilized very effectively in detecting and parallelizing complex
code. To demonstrate this, we have taken two complex benchmarks -
435.gromacs and 437.leslie3d from the SPECfp CPU2006 suite and show how
they can be parallelized effectively using DProf as an assist. To the
best of our knowledge none of the existing parallelizing compilers can
detect and parallelize all the instances reported in this work. We find
that by using DProf we are able to parallelize these benchmarks very
effectively for IBM P5+ and P6 multi-core systems. We have parallelized
the benchmarks using OpenMP leading to speedups up to 2.6x. Also, due
to the detailed reporting by DProf, we could cut down on our
parallelization development effort significantly by concentrating on
portions of the code that require attention. DProf can also be used to
identify applications, where applying parallelization may lead to
regression in performance. This allows the developers to discard
applications or parts of it quickly, which may not lead to performance
improvements when deployed on multi-cores. Thus, a data dependence
profiler like DProf can act as an excellent assist mechanism to move
applications to multi-cores in an effective way.

#e1#

#s1#
#abstract1357.txt#


The Gozer workflow system is a production workflow authoring and
execution platform that was developed at RiskMetrics Group. It provides
a high-level language and supporting libraries for implementing local
and distributed parallel processes. Gozer was developed with an
emphasis on distributed processing environments in which workflows may
execute for hours or even days. Key features of Gozer include: implicit
parallelization that exploits both local and distributed parallel
resources; survivability of system faults/shutdowns without losing
state; automatic distributed process migration; and implicit resource
management and control. The Gozer language is a dialect of Lisp, and
the Gozer system is implemented on a service-oriented architecture.

#e1#

#s1#
#abstract1358.txt#


Simulation is a "simple" method to experimentally evaluate the behavior
of algorithms designed for parallel and distributed platforms.
Moreover, the reliability of the evaluation strongly depends on the
models used inside the simulator. This paper is devoted to the study
and the evaluation of the behavior of the SimGrid simulator through a
comparison between a simulated and a real execution of the same target
application (i.e. heat propagation). Our target platforms are
heterogeneous and their characteristics may dynamically vary during the
execution. The obtained results (on a set of various platforms) show
that the behavior observed when using the simulated platform is very
close to the one obtained on the real one.

#e1#

#s1#
#abstract1359.txt#


Nowadays, commodity computers are complex heterogeneous systems that
provide a huge amount of computational power. However, to take
advantage of this power we have to orchestrate the use of processing
units with different characteristics. Such distributed memory systems
make use of relatively slow interconnection networks, such as system
buses. Therefore, most of the time we only individually take advantage
of the central processing unit (CPU) or processing accelerators, which
are simpler homogeneous subsystems. In this paper we propose a
collaborative execution environment for exploiting data parallelism in
a heterogeneous system. It is shown that this environment can be
applied to program both CPU and graphics processing units (GPUs) to
collaboratively compute matrix multiplication and fast Fourier
transform (FFT). Experimental results show that significant performance
benefits are achieved when both CPU and GPU are used.

#e1#

#s1#
#abstract1360.txt#


This paper studies the difference in computational power between the
mesh-connected parallel computers equipped with dynamically
reconfigurable bus systems and those with static ones. The mesh with
separable buses (MSB) is the mesh-connected computer with dynamically
reconfigurable row/column buses. The broadcasting buses of the MSB can
be dynamically sectioned into smaller bus segments by program control.
We examine the impact of reconfigurable capability on the computational
power of the MSB model, and investigate how computing power of the MSB
decreases when we deprive the MSB of its reconfigurability. We show
that any single step of the MSB of size n * n can be simulated in O
(log n) time by the MSB without its reconfigurable function, which
means that the MSB of size n * n can work with O (log n) step slowdown
even if its dynamic reconfigurable function is disabled.

#e1#

#s1#
#abstract1361.txt#


The computational power provided by the massive parallelism of modern
graphics processing units (GPUs) has moved increasingly into focus over
the past few years. In particular, general purpose computing on GPUs
(GPGPU) is attracting attention among researchers and practitioners
alike. Yet GPGPU research is still in its infancy, and a major
challenge is to rearrange existing algorithms so as to obtain a
significant performance gain from the execution on a GPU. In this
paper, we address this challenge by presenting an efficient GPU
implementation of a very popular algorithm for linear programming, the
revised simplex method. We describe how to carry out the steps of the
revised simplex method to take full advantage of the parallel
processing capabilities of a GPU. Our experiments demonstrate
considerable speedup over a widely used CPU implementation, thus
underlining the tremendous potential of GPGPU.

#e1#

#s1#
#abstract1362.txt#


The discrete wavelet transform (DWT) is a powerful signal processing
technique used in the JPEG 2000 image compression standard. The
multi-resolution sub-band encoding provided by DWT allows for higher
compression ratios, avoids blocking artifacts and enables progressive
transmission of images. However, these advantages come at the expense
of additional computational complexity. Achieving real-time or
interactive compression/de-compression speeds, therefore, requires a
fast implementation of DWT that leverages emerging parallel hardware
systems. In this paper, we develop an optimized parallel implementation
of the lifting-based DWT algorithm using the recently proposed Open
Computing Language (OpenCL). OpenCL is a standard for cross-platform
parallel programming of heterogeneous systems comprising of multi-core
CPUs, GPUs and other accelerators. We explore the potential of OpenCL
in accelerating the DWT computation and analyze the programmability,
portability and performance aspects of this language. Our experimental
analysis is done using NVIDIA's and AMD's drivers that support OpenCL.

#e1#

#s1#
#abstract1363.txt#


Mutual-Information-Based Registration (MIBR) is an image registration
method that maps points from one image to another. It has been widely
used in medical image processing applications. However, MIBR is a very
compute-intensive task, and fast processing speed is often required in
medical diagnosis. Nowadays, with the multi-core processor becoming the
mainstream, MIBR can be accelerated by fully utilizing the computing
power of available multi-core processors. In this paper, we propose a
parallel MIBR algorithm and present some optimization techniques to
improve the implementation's performance. The result shows our
optimized implementation can register a pair of 512 * 512 * 30 3D
images in one second on an 8-core system, which meets the real-time
processing requirement. We also conduct a detailed scalability and
memory performance analysis on the multi-core system. The analysis
helps us to identify the causes of bottlenecks, and make suggestion for
future improvement on large-scale multi-core systems.

#e1#

#s1#
#abstract1364.txt#


Program analysis supporting software development is often part of
edit-compile-cycles, and precise program analysis is time consuming.
With the availability of parallel processing power on desktop
computers, parallelization is a way to speed up program analysis. This
requires a parallel data-flow analysis with sufficient work for each
processing unit. The present paper suggests such an approach for
object-oriented programs analyzing the target methods of polymorphic
calls in parallel. With carefully selected thresholds guaranteeing
sufficient work for the parallel threads and only little redundancy
between them, this approach achieves a maximum speed-up of 5 (average
1.78) on 8 cores for the benchmark programs.

#e1#

#s1#
#abstract1365.txt#


Graphics processing units provide a large computational power at a very
low price which position them as an ubiquitous accelerator. General
purpose programming on the graphics processing units (GPGPU) is best
suited for regular data parallel algorithms. They are not directly
amenable for algorithms which have irregular data access patterns such
as list ranking, and finding the connected components of a graph, and
the like. In this work, we present a GPU-optimized implementation for
finding the connected components of a given graph. Our implementation
tries to minimize the impact of irregularity, both at the data level
and functional level. Our implementation achieves a speed up of 9 to 12
times over the best sequential CPU implementation. For instance, our
implementation finds connected components of a graph of 10 million
nodes and 60 million edges in about 500 milliseconds on a GPU, given a
random edge list. We also draw interesting observations on why PRAM
algorithms, such as the Shiloach-Vishkin algorithm may not be a good
fit for the GPU and how they should be modified.

#e1#

#s1#
#abstract1366.txt#


In studying the scalability of the Scalasca performance analysis
toolset to several hundred thousand MPI processes on IBM Blue Gene/P,
we investigated a progressive execution performance deterioration of
the well-known ASCI Sweep3D compact application. Scalasca runtime
summarization analysis quantified MPI communication time that
correlated wth computational imbalance, and automated trace analysis
confirmed growing amounts of MPI waiting times. Further
instrumentation, measurement and analyses pinpointed a conditional
section of highly imbalanced computation which amplified waiting times
inherent in the associated wavefront communication that seriously
degraded overall execution efficiency at very large scales. By
employing effective data collation, management and graphical
presentation, Scalasca was thereby able to demonstrate performance
measurements and analyses with 294,912 processes for the first time.

#e1#

#s1#
#abstract1367.txt#


New algorithms and optimization techniques are needed to balance the
accelerating trend towards bandwidth-starved multicore chips. It is
well known that the performance of stencil codes can be improved by
temporal blocking, lessening the pressure on the memory interface. We
introduce a new pipelined approach that makes explicit use of shared
caches in multicore environments and minimizes synchronization and
boundary overhead. For clusters of shared-memory nodes we demonstrate
how temporal blocking can be employed successfully in a hybrid
shared/distributed-memory environment.

#e1#

#s1#
#abstract1368.txt#


The increasing core-count on current and future processors is posing
critical challenges to the memory subsystem to efficiently handle
concurrent memory requests. The current trend is to increase the number
of memory channels available to the processor's memory controller. In
this paper we investigate the effectiveness of this approach on the
performance of parallel scientific applications. Specifically, we
explore the trade-off between employing multiple memory channels per
memory controller and the use of multiple memory controllers.
Experiments conducted on two current state-of-the-art multicore
processors, a 6-core AMD Istanbul and a 4-core Intel Nehalem-EP, for a
wide range of production applications shows that there is a diminishing
return when increasing the number of memory channels per memory
controller. In addition, we show that this performance degradation can
be efficiently addressed by increasing the ratio of memory controllers
to channels while keeping the number of memory channels constant.
Significant performance improvements can be achieved in this scheme, up
to 28%, in the case of using two memory controllers each with one
channel compared with one controller with two memory channels.

#e1#

#s1#
#abstract1369.txt#


A large part of today's most popular applications are data-intensive.
Whether they are scientific applications or Internet services, the data
volume they process is continuously growing. Two main aspects arise
when trying to accomodate the size of the data: processing the
computation in a manner that is efficient both in terms of resources
and time, and providing storage capable to deal with the requirements
of data-intensive applications. Since the input data is large, the
computation, which is, in most cases straightforward, is distributed
across hundreds or thousands of machines; thus, the application is
split into tasks that run in parallel on different machines, tasks that
will need to access the data in a highly concurrent manner.

#e1#

#s1#
#abstract1370.txt#


As the rate, scale and variety of data increases in complexity, the
need for flexible applications that can crunch huge amounts of
heterogeneous data fast and cost-effective is of utmost importance.
Such applications are data-intensive: in a typical scenario, they
continuously acquire massive datasets (e.g. by crawling the Web or
analyzing access logs) while performing computations over these
changing datasets (e.g. building up-to-date search indexes). In order
to achieve scalability and performance, data acquisitions and
computations need to be distributed at large scale in infrastructures
comprising hundreds and thousands of machines. As these applications
focus on data rather then on computation, a heavy burden is put on the
storage service employed to handle data management, because it must
efficiently deal with massively parallel data accesses. In order to
achieve this, a series of issues need to be address properly: scalable
aggregation of storage space from the participating nodes with minimal
overhead, the ability to store huge data objects, efficient fine-grain
access to data subsets, high throughput even under heavy access
concurrency, versioning, as well as fault tolerance and a high quality
of service for access throughput. This paper introduces BlobSeer, an
efficient distributed data management service that addresses the issues
presented above. In BlobSeer, long sequences of bytes representing
unstructured data are called blobs (Binary Large OBject).

#e1#

#s1#
#abstract1371.txt#


Though current general-purpose processors have several small CPU cores
as opposed to a single more complex core, many algorithms and
applications are inherently sequential and so hard to explicitly
parallelize. Cores designed to handle these problems may exhibit deeper
pipelines and wider fetch widths to exploit instruction-level
parallelism via out-of-order execution. As these parameters increase,
so does the amount of instructions fetched along an incorrect path when
a branch is mispredicted. Some instructions are fetched regardless of
the direction of a branch. In current conventional CPUs, these
instructions are always squashed upon branch misprediction and are
fetched again shortly thereafter. Recent research efforts explore
lessening the effect of branch mispredictions by retaining these
instructions when squashing or fetching them in advance when
encountering a branch that is difficult to predict. Though these
control independent processors are meant to lessen the damage of
misprediction, an inherent side-effect of fetching out of order, branch
weakening, reduces realized speedup and is in part responsible for
lowering potential speedup. This study formally defines and works
towards identifying the causes of branch weakening. The overall goal of
the research is to determine how much weakening is avoidable and
develop techniques to help reduce weakening in control independent
processors.

#e1#

#s1#
#abstract1372.txt#


MPI (Message Passing Interface) has been successfully used in the high
performance computing community for years and is the dominant
programming model. Current implementations of MPI are coarse-grained,
with a single MPI process per processor, however, there is nothing in
the MPI specification precluding a finer-grain interpretation of the
standard. We have implemented Fine-grain MPI (FG-MPI), a system that
allows execution of hundreds and thousands of MPI processes on-chip or
communicating between chips inside a cluster. FG-MPI uses fibers
(coroutines) to support multiple MPI processes inside an operating
system process. These are fullfledged MPI processes each with their own
MPI rank. We have implemented a fine-grain version of MPICH2 middleware
that uses the Nemesis communication subsystem for intranode and
internode communication. We present experimental results for a
real-world application that uses thousands of MPI processes and compare
its performance with the following fine-grain multicore languages:
Erlang, Haskell, Occam-pi and POSIX threads. Our results show that
FG-MPI scales well and outperforms many of these other programming
languages used for parallel programming on multicore systems while
retaining MPI's intranode and internode communication abilities.

#e1#

#s1#
#abstract1373.txt#


Recently, Graphical Processing Units (GPUs) have become increasingly
more capable and well-suited to general purpose applications. As a
result of the GPUs high degree of parallelism and computational power,
there has been a great deal of interest directed toward the platform
for parallel application development. Much of the focus, however, has
been on very regular applications that exhibit a high degree of data
parallelism, as these applications map well to the GPU. Irregular
applications, such as the Breadth First Search discussed in this paper,
have not been as extensively studied and are more difficult to
implement in an efficient fashion on the GPU. We will present both an
implementation of the Breadth First Search algorithm as well as that of
a Matrix Parenthesization algorithm. These pair of algorithms showcase
similar synchronization behavior when implemented on a GPU using CUDA,
enabling a more direct comparison between them. The results obtained
can be used to showcase some of the synchronization issues present with
irregular algorithms on the GPU.

#e1#

#s1#
#abstract1374.txt#


Current Graphics Processing Unit (GPU) presents large potentials in
speeding up computationally intensive data parallel applications over
traditional parallelization approaches since there are much more
hardware threads inside GPUs than the computational cores available to
common CPU threads. NVIDIA developed a generic GPU programming
platform, CUDA, which allows programmers to utilize GPU through C
programming language and parallelize applications in a similar way as
in traditional multithreading approach. However, not all applications
are suitable for this new platform. Only computationally intensive
applications without strong dependency are good candidates. Although
Advanced Encryption Standard (AES) does not belong to this group due to
the light workload in its efficient implementation, this paper proposed
an approach to arrange data in different GPU memory spaces properly,
overcoming the extra communication delay, and still turning GPU into an
effective accelerator. Experimental results have demonstrated its
effectiveness by performance gains and proved that GPU can be used to
accelerate more types of applications.

#e1#

#s1#
#abstract1375.txt#


As multi-cores arrive for mainstream desktop systems, developers must
invest the effort to parallelize their applications. We present
Parallel Task (short ParaTask), a solution to assist the
parallelization of object-oriented applications, with the unique
feature of including support for the parallelization of graphical user
interface (GUI) applications. In the simple, but common, cases
concurrency is introduced with a single keyword. Due to the wide
variety of parallelization needs, ParaTask integrates different task
types into the same model, provides intuitive support for dependence
handling, non-blocking notification, interim progress notification and
exception handling in an asynchronous environment as well as supporting
a pluggable task scheduling runtime (currently work-sharing,
work-stealing and a combination of the two are supported). The
performance is compared to traditional Java parallelization approaches
using a variety of different workloads.

#e1#

#s1#
#abstract1376.txt#


Quantum chemistry applications such as the General Atomic and Molecular
Electronic Structure System (GAMESS) that can execute on a complex
peta-scale parallel computing environment has a large number of input
parameters that affect the overall performance. The application
characteristics vary according to the input parameters. This is due to
the difference in the usage of resources like network bandwidth, I/O
and main memory, according to the input parameters. Effective execution
of applications in a parallel computing environment that share such
resources require some sort of adaptive mechanism to enable efficient
usage of these resources. In our previous work, we have integrated
GAMESS with an adaptive middleware NICAN (Network Information Conveyer
and Application Notification) for dynamic adaptations during heavy load
conditions that modify execution of GAMESS computations on a
per-iteration basis. This leads to better application performance. In
this research, we have expanded the structure of NICAN in order to
include other input parameters based on which application performance
can be controlled. The application performance has been analyzed on
different architectures and a tuning strategy has been identified. A
generic database framework has been incorporated in the existing NICAN
mechanism so as to aid this tuning strategy.

#e1#

#s1#
#abstract1377.txt#


In this paper, we present an approach for a patch-based adaptive mesh
refinement (AMR) for multi-physics simulations. The approach consists
of clustering, symmetry preserving, mesh continuity, flux correction,
communications, management of patches, and dynamical load balance.
Among the special features of this patch-based AMR are symmetry
preserving, efficiency of refinement, special implementation of flux
correction, and patch management in parallel computing environments.
Here, higher efficiency of refinement means less unnecessarily refined
cells for a given set of cells to be refined. To demonstrate the
capability of the AMR framework, hydrodynamics simulations with many
levels of refinement are shown in both two- and three-dimensions.

#e1#

#s1#
#abstract1378.txt#


This paper focuses on parallelization of the classic static timing
analysis (STA) algorithm for verifying timing characteristics of
digital integrated circuits. Given ever-increasing circuit
complexities, including the need to analyze circuits with billions of
transistors, across potentially thousands of process corners, with
accuracy tolerances down to the picosecond range, sequential execution
of STA algorithms is quickly becoming a bottleneck to the overall chip
design closure process. A message passing based parallel processing
technique for performing STA leveraging an IBM Blue Gene/L
supercomputing platform is presented. Results are collected for a small
industrial 65 nm benchmarking design, where the algorithm demonstrates
speedup of nearly 39 times on 64 processors and a peak of 119 times
(without partitioning costs, speedup is 263 times) on 1024 processors.
With an idealized synthetic circuit, the algorithm demonstrated 259
times speedup, 925 times speedup without partitioning overhead, on 1024
processors. To the best of our knowledge, this is the first result
demonstrating scalable STA on the IBM Blue Gene.

#e1#

#s1#
#abstract1379.txt#


Fast Block Matching (FBM) algorithms for video compression are well
suited for acceleration using parallel data-path architecture on Field
Programmable Gate Arrays (FPGAs). However, designing an efficient
on-chip memory subsystem to provide the required throughput to this
parallel data-path architecture is a complex problem. This paper
proposes a memory architecture template that is explored using a
Bounded Set algorithm to design efficient on-chip memory subsystems for
FBM algorithms. The resulting memory subsystems are compared with three
existing memory subsystems. Results show that our memory subsystems can
provide full parallelism in majority of test cases and can process
integer pixels of a 1080 p video sequence up to a rate of 275 frames
per second.

#e1#

#s1#
#abstract1380.txt#


Limits of instruction level parallelism and the higher transistor
density sustain the increasing need for multiprocessor systems: they
are rapidly taking over both general purpose and embedded processor
domains. Nowadays, since these processors must handle a wide range of
different application classes, there is no consensus over which are the
best hardware solutions to exploit the best of ILP and TLP together.
Current multiprocessing systems are composed either of many homogeneous
and simple cores, or of complex superscalar SMT processing elements. In
this work, we have expanded a reconfigurable architecture to be used in
a multiprocessing scenario, showing the need for an adaptable ILP
exploitation even in TLP architectures. We have successfully coupled a
dynamic reconfigurable system to a SPARC-based multiprocessor, and
obtained performance gains of up to 40% even for applications that show
a great level of parallelism at thread level, demonstrating the need
for an adaptable ILP exploration.

#e1#

#s1#
#abstract1381.txt#


The k-nearest neighbor (k-NN) is a popular non-parametric benchmark
classification algorithm to which new classifiers are usually compared.
It is used in numerous applications, some of which may involve
thousands of data vectors in a possibly very high dimensional feature
space. For real-time classification a hardware implementation of the
algorithm can deliver high performance gains by exploiting parallel
processing and block pipelining. We present two different linear array
architectures that have been described as soft parameterized IP cores
in VHDL. The IP cores are used to synthesize and evaluate a variety of
array architectures for a different k-NN problem instances and Xilinx
FPGAs. It is shown that we can solve efficiently, using a medium size
FPGA device, very large size classification problems, with thousands of
reference data vectors or vector dimensions, while achieving very high
throughput. To the best of our knowledge, this is the first effort to
design flexible IP cores for the FPGA implementation of the widely used
k-NN classifier.

#e1#

#s1#
#abstract1382.txt#


This work presents a scalability analysis of embarrassingly parallel
applications running on cluster and multi-cluster machines. Several
applications can be included in this category. Examples are
Bag-of-tasks (BoT) applications and some classes of online web
services, such as index processing in online web search. The analysis
presented here is divided in two parts: first, the impact of front end
topology on scalability is assessed through a lower bound analysis. In
a second step several task mapping strategies are compared from the
scalability standpoint.

#e1#

#s1#
#abstract1383.txt#


Parallel NFS (pNFS) is touted as an emergent standard protocol for
parallel I/O access in various storage environments. Several pNFS
prototypes have been implemented for initial validation and protocol
examination. Previous efforts have focused on realizing the pNFS
protocol to expose the best bandwidth potential from underlying file
and storage systems. In this presentation, we provide an initial
characterization of two pNFS prototype implementations, lpNFS (a
Lustre-based parallel NFS implementation) and spNFS (another reference
implementation from Network Appliance, Inc.). We show that both lpNFS
and spNFS can faithfully achieve the primary goal of pNFS, i.e.,
aggregating I/O bandwidth from many storage servers. However, they both
face the challenge of scalable metadata management. Particularly, the
throughput of sp-NFS metadata operations degrades significanlty with an
increasing number of data servers. Even for the better-performing
lpNFS, we discuss its architecture and propose a direct I/O request
flow protocol to improve its performance.

#e1#

#s1#
#abstract1384.txt#


The following topics are dealt with: parallel processing; distributed
processing; reconfigurable architectures; high-level parallel
programming; nature inspired distributed scientific computing; high
performance computational biology; communication architecture;
power-aware computing; high performance grid computing; system
management techniques; system management processes; system management
services; parallel computing; engineering computing; ubiquitous
computing; optimisation; network-centric systems; peer-to-peer systems;
finance computing and multi-threaded architectures.

#e1#

#s1#
#abstract1385.txt#


We present a Graphics Processing Unit (GPU) parallelization of the
computation of the price of cross-currency interest rate derivatives
via a Partial Differential Equation (PDE) approach. In particular, we
focus on the GPU-based parallel computation of the price of long-dated
foreign exchange interest rate hybrids, namely Power Reverse Dual
Currency (PRDC) swaps with Bermudan cancelable features. We consider a
three-factor pricing model with foreign exchange skew which results in
a time-dependent parabolic PDE in three spatial dimensions. Finite
difference methods on uniform grids are used for the spatial
discretization of the PDE, and the Alternating Direction Implicit (ADI)
technique is employed for the time discretization. We then exploit the
parallel architectural features of GPUs together with the Compute
Unified Device Architecture (CUDA) framework to design and implement an
efficient parallel algorithm for pricing PRDC swaps. Over each period
of the tenor structure, we divide the pricing of a Bermudan cancelable
PRDC swap into two independent pricing subproblems, each of which can
efficiently be solved on a GPU via a parallelization of the ADI scheme
at each timestep. Using this approach on two NVIDIA Tesla C870 GPUs of
an NVIDIA 4-GPU Tesla S870 to price a Bermudan cancelable PRDC swap
having a 30 year maturity and annual exchange of fund flows, we have
achieved an asymptotic speedup by a factor of 44 relative to a single
thread on a 2.0GHz Xeon processor.

#e1#

#s1#
#abstract1386.txt#


We propose a new parallel asynchronous cellular genetic algorithm for
multi-core processors. The algorithm is applied to the scheduling of
independent tasks in a grid. Finding such optimal schedules is in
general an NP-hard problem, to which evolutionary algorithms can find
near-optimal solutions. We analyze the parallelism of the algorithm, as
well as different recombination and new local search operators. The
proposed algorithm improves previous schedules on benchmark problems.
The parallelism of this algorithm suits it to bigger problem instances.

#e1#

#s1#
#abstract1387.txt#


This paper deals with the problem of identifying faulty nodes (or
units) in diagnosable distributed and parallel systems under the PMC
model. In this model, each unit is tested by a subset of the other
units, and it is assumed that, at most, a bounded subset of these units
is permanently faulty. When performing testing, faulty units can
incorrectly claim that fault-free units are faulty or that faulty units
are fault-free. Since the introduction of the PMC model, significant
progress has been made in both theory and practice associated with the
original model and its offshoots. Nevertheless, this problem of
efficiently identifying the set of faulty units of a diagnosable system
remained an outstanding research issue. In this paper, we describe a
new neural-network-based diagnosis algorithm, which exploits the
off-line learning phase of artificial neural network to speed up the
diagnosis algorithm. The novel approach has been implemented and
evaluated using randomly generated diagnosable systems. The simulation
results showed that the new neural-network-based fault identification
approach constitutes an addition to existing diagnosis algorithms.
Extreme faulty situations, where the number of faults is around the
bound t, and large diagnosable systems have been also experimented to
show the efficiency of the new neural-network-based diagnosis algorithm.

#e1#

#s1#
#abstract1388.txt#


The increasing availability of multi-core and multiprocessor
architectures provides new opportunities for improving the performance
of many computer simulations. Markov Chain Monte Carlo (MCMC)
simulations are widely used for approximate counting problems, Bayesian
inference and as a means for estimating very high-dimensional
integrals. As such MCMC has had a wide variety of applications in
fields including computational biology and physics, financial
econometrics, machine learning and image processing. One method for
improving the performance of Markov Chain Monte Carlo simulations is to
use SMP machines to perform 'speculative moves', reducing the runtime
whilst producing statistically identical results to conventional
sequential implementations. In this paper we examine the circumstances
under which the original speculative moves method performs poorly, and
consider how some of the situations can be addressed by refining the
implementation. We extend the technique to perform Markov Chains
speculatively, expanding the range of algorithms that maybe be
accelerated by speculative execution to those with non-uniform move
processing times. By simulating program runs we can predict the
theoretical reduction in runtime that may be achieved by this
technique. We compare how efficiently different architectures perform
in using this method, and present experiments that demonstrate a
runtime reduction of up to 35-42% where using conventional speculative
moves would result in execution as slow, if not slower, than sequential
processing.

#e1#

#s1#
#abstract1389.txt#


There is building interest in using FPGAs as accelerators for
high-performance computing, but existing systems for programming them
are so far inadequate. In this paper we propose a soft processor
programming model and architecture inspired by graphics processing
units (GPUs) that are well-matched to the strengths of FPGAs, namely
highly-parallel and pipelinable computation. In particular, our soft
processor architecture exploits multithreading and vector operations to
supply a floating-point pipeline of 64 stages via hardware support for
up to 256 concurrent thread contexts. The key new contributions of our
architecture are mechanisms for managing threads and register files
that maximize data-level and instruction-level parallelism while
overcoming the challenges of port limitations of FPGA block memories,
as well as memory and pipeline latency. Through simulation of a system
that (i) supports AMD's CTM r5xx GPU ISA , and (ii) is realizable on an
XtremeData XD1000 FPGA-based accelerator system, we demonstrate that
our soft processor can achieve 100% utilization of the deeply-pipelined
floating-point datapath.

#e1#

#s1#
#abstract1390.txt#


We introduce a high performance parallelization to the PSTD solution of
Maxwell equations by employing the fast Fourier transform on local
Fourier basis. Meanwhile a reformatted derivative operator allows the
adoption of a staggered-grid such as the Yee lattice in PSTD, which can
overcome the numerical errors in a collocated-grid when spatial
discontinuities are present. The accuracy and capability of our method
are confirmed by two analytical models. In two applications to surface
tissue optics, an ultra wide coherent backscattering cone from the
surface layer is found, and the penetration depth of polarization
gating identified. Our development prepares a tool for investigating
the optical properties of surface tissue structures.

#e1#

#s1#
#abstract1391.txt#


In recent years, researchers have found that some XOR erasure codes
lead to higher performance and better throughput in fault-tolerant
distributed data storage applications. However, little consideration
has been given to the advantages of parallel processing or hardware
implementations taking advantage of the emergence of multi-core
processors. This paper presents an efficient horizontal MDS-like
(Maximum Distance Separable) RAID-6 scheme, called EEO, which
significantly improves the performance of the decoding procedure in
parallel implementations with little storage overhead. We show that EEO
is the fastest and most efficient double disk failure recovering
algorithm in RAID-6 at the cost of only two more parity symbols. In
practice, it is very useful for application where high decoding
throughput is desired.

#e1#

#s1#
#abstract1395.txt#


Traditional System on Chip (SOC) designs offer integrated solutions to
exigent design tribulations in areas which necessitate outsized
computation and restriction in certain area. But the performance of
these has been sluggish due to the restriction of the common bus
architecture espoused by these systems and thereby low processing
speeds. This has been the main stumbling block for scalability in terms
of computation and enhancement in its performance. With the advancement
in semi conductor devices and fabrication technology, it is possible to
pack more logic in smaller area of silicon. But the implementation of
these mega functional modules using common bus architecture, parallel
bus architecture, pipelining are becoming ineffective and posing a
bottleneck in terms of performance and throughput in this billion
transistor era. As an elucidation for this problem, Network on chip is
being adopted in this paper as the core bus architecture across
different spectrum of SOCs. Our approach presents a supple design using
FPGA based system. Hence, it is a very flexible network design that
will accommodate to various needs.

#e1#

#s1#
#abstract1396.txt#


In today's world application of battery powered analog and mixed mode
electronic device requires designing analog circuit to operate at low
voltage levels. There are many issues which are involved in
implementing low voltage circuits such as reduced noise immunity,
greater delay and poor linearity. There is a need to use certain MOSFET
techniques so that the MOSFET can be used even in sub-threshold region
without any significant change in its performance. Certain techniques
have been put forward corresponding to low voltage analog circuits
using CMOS Technology to be used in variety of applications.
Accordingly we propose to design a 10 bits 30Msample/s CMOS Analog to
Digital Converter (ADC) using a 1.5-bits/stage pipeline architecture
for high speed signal processing to be used in video related
applications.

#e1#

#s1#
#abstract1397.txt#


To index, search, browse and retrieve relevant material, indexes
describing the video content are required. Here, a new and fast
strategy which allows detecting abrupt and gradual transitions is
proposed. A pixel-based analysis is applied to detect abrupt
transitions and, in parallel, an edge-based analysis is used to detect
gradual transitions. Both analysis are reinforced with a motion
analysis in a second step, which significantly simplifies the threshold
selection problem while preserving the computational requirements. The
main advantage of the proposed system is its ability to work in real
time and the experimental results show high recall and precision values.

#e1#

#s1#
#abstract1398.txt#


K-Means is a clustering algorithm that is widely applied in many
fields, including pattern classification and multimedia analysis. Due
to real-time requirements and computational-cost constraints in
embedded systems, it is necessary to accelerate K-Means algorithm by
hardware implementations in SoC environments, where the bandwidth of
the system bus is strictly limited. In this paper, a bandwidth adaptive
hardware architecture of K-Means clustering is proposed. Experiments
show that the proposed hardware can be used in applications such as
image segmentation, and it has the maximum clock speed 400-MHz and
440-K gate count with TSMC 90-nm technology. Moreover, the throughput
of the proposed hardware reaches 16 dimension/cycle, and it can deal
with feature vectors with different dimensions using five parallel
modes to utilize the input bandwidth efficiently.

#e1#

#s1#
#abstract1399.txt#


This work presents mathematical models and collision-free exchange
rules for a parallel interleaver, using which it develops an optimized
memory address remapping (OPMM) scheme that enables a classic
interleaver to be exchanged for a parallel interleaver readily and
efficiently. Both analytic and experimental results demonstrate that
the rate of annealing achieved using the OPMM approach is much faster
than that achieved using the traditional memory address remapping (MM)
method.

#e1#

#s1#
#abstract1400.txt#


Variable block size motion estimation (VBSME) is one of several
contributors to H.264/AVC's excellent coding efficiency. However, its
high computational complexity and huge memory traffic make deign
difficult. In this paper, we propose a memory-efficient and highly
parallel VLSI architecture for full search VBSME (FSVBSME). Our
architecture consists of 16 2-D arrays each consists of 16 * 16
processing elements (PEs). Four arrays form a group to match in
parallel four reference blocks against one current block. Four groups
perform block matching for four current blocks in a pipelined fashion.
Taking advantage of overlapping among multiple reference blocks of a
current block and between search windows of adjacent current blocks, we
propose a novel data reuse scheme to reduce memory access. Compared
with the popular Level C data reuse scheme, our approach can save 98%
of on-chip memory access with only 25% of local memory overhead.
Synthesized into a TSMC 180-nm CMOS cell library, our design is capable
of processing 1920 * 1088 30 fps video when running at 130 MHz. The
architecture is scalable for wider search range, multiple reference
frames and pixel truncation as well as down sampling. We suggest a
criterion called design efficiency for comparing different works. It
shows that the proposed design is 72% more efficient than the best
design to date.

#e1#

#s1#
#abstract1401.txt#


A high frequency ultrasound-coupled fluorescence tomography system,
primarily designed for imaging of protoporphyrin IX production in skin
tumors in vivo, is demonstrated for the first time. The design couples
fiber-based spectral sampling of the protoporphyrin IX fluorescence
emission with high frequency ultrasound imaging, allowing thin-layer
fluorescence intensities to be quantified. The system measurements are
obtained by serial illumination of four linear source locations, with
parallel detection at each of five interspersed detection locations,
providing 20 overlapping measures of subsurface fluorescence from both
superficial and deep locations in the ultrasound field. Tissue layers
are defined from the segmented ultrasound images and diffusion theory
used to estimate the fluorescence in these layers. The system
calibration is presented with simulation and phantom validation of the
system in multilayer regions. Pilot in-vivo data are also presented,
showing recovery of subcutaneous tumor tissue values of protoporphyrin
IX in a subcutaneous U251 tumor, which has less fluorescence than the
skin.

#e1#

#s1#
#abstract1402.txt#


A new, freely available third party MATLAB toolbox for the simulation
and reconstruction of photoacoustic wave fields is described. The
toolbox, named k-Wave, is designed to make realistic photoacoustic
modeling simple and fast. The forward simulations are based on a
k-space pseudo-spectral time domain solution to coupled first-order
acoustic equations for homogeneous or heterogeneous media in one, two,
and three dimensions. The simulation functions can additionally be used
as a flexible time reversal image reconstruction algorithm for an
arbitrarily shaped measurement surface. A one-step image reconstruction
algorithm for a planar detector geometry based on the fast Fourier
transform (FFT) is also included. The architecture and use of the
toolbox are described, and several novel modeling examples are given.
First, the use of data interpolation is shown to considerably improve
time reversal reconstructions when the measurement surface has only a
sparse array of detector points. Second, by comparison with one-step,
FFT-based reconstruction, time reversal is shown to be sufficiently
general that it can also be used for finite-sized planar measurement
surfaces. Last, the optimization of computational speed is demonstrated
through parallel execution using a graphics processing unit.

#e1#

#s1#
#abstract1403.txt#


The data acquisition speed in photoacoustic computed tomography (PACT)
is limited by the laser repetition rate and the number of parallel
ultrasound detecting channels. Reconstructing an image with fewer
measurements can effectively accelerate the data acquisition and reduce
the system cost. We adapt compressed sensing (CS) for the
reconstruction in PACT. CS-based PACT is implemented as a nonlinear
conjugate gradient descent algorithm and tested with both phantom and
in vivo experiments.

#e1#

#s1#
#abstract1404.txt#


In the frame of biological threat, security systems require label free
biochips for rapid detection. Biosensors enable to detect biological
interactions, between probes localized at the surface of a chip, and
targets present in the sample solution. Here, we present an optical
transduction, enabling 2D imaging, and consequently parallel detection
of several reactions. It is based on the absorption of biological
molecules in the UV domain. Thus, it is based on an intrinsic property
of biological molecules and does not require any labelling of the
biological molecules. DNA and proteins absorb UV light at 260 and 280
nm respectively. Sensitivity is a major requirement of biosensing
devices. Configurations leading to enhancement of the interaction
between light and biological molecules are of interest. For a better
sensitivity, resonant grating structures are then studied. They enable
to confine the electric field close to the biological layer. Imaging of
resonant grating is not largely studied, even for visible wavelengths,
but it results in good sensitivity. The protein used in this study is
the methionyl-tRNA synthetase. Its absorption is representative of
protein absorption, and it can then serve as a model for immunological
detection. The best experimental contrast due to a monolayer of
proteins is 40%. With data processing currently employed for biochip
imaging: average on several acquisitions and on all the pixels imaging
the biological spots, the device is able to detect a surface density of
proteins in the 10 pg/mm range.

#e1#

#s1#
#abstract1405.txt#


A new design of multipliers for GF(2/sup m/) based on combination of
bit-serial and bit-parallel schemes with low complexity is proposed.
Using pipeline architecture, the scheme yields significantly lower
latency compared to known bit-parallel multipliers for GF(2/sup m/).

#e1#

#s1#
#abstract1406.txt#


This paper addresses decoder design for nonbinary quasicyclic
low-density parity-check (QC-LDPC) codes. First, a novel decoding
algorithm is proposed to eliminate the multiplications over Galois
field for check node processing. Then, a partially parallel
architecture for check node processing units and an optimized
architecture for variable node processing units are developed based on
the new decoding algorithm. Thereafter, an efficient decoder structure
dedicated to a promising class of high-performance nonbinary QC-LDPC
codes is presented for the first time. Moreover, an ASIC implementation
for a (620, 310) nonbinary QC-LDPC code decoder over GF(32) is designed
to demonstrate the efficiency of the presented techniques.

#e1#

#s1#
#abstract1407.txt#


A report on the use of programmable gate-ion sensitive field effect
transistors (PG-ISFETs) to form DNA-logic gates for reaction monitoring
is presented. A PG-ISFET sensor is integrated with standard MOSFETs to
form a chemical-inverter that switches when a chemical reaction reaches
a certain threshold. By employing feedback to the electrical gate rapid
switching is achieved for changes in pH as low as 0.1 units, overcoming
limitations of voltage division owing to passivation capacitance of
CMOS based ISFETs. Using this, DNA-logic arrays can be synthesised to
carry out highly parallel chemical processing functions for
applications such as DNA sequencing.

#e1#

#s1#
#abstract1408.txt#


This article contains a short analysis of applying three metaheuristic
local search algorithms to solve the problem of allocating
two-dimensional tasks on a two-dimensional processor mesh in a period
of time. The primary goal is to maximize the level of mesh utilization.
To achieve this task we adapted three algorithms: Tabu Search,
Simulated Annealing and Random Search, as well as created a helper
algorithm Dumb Fit and adapted another helper algorithm First Fit. To
measure the algorithms' efficiency we introduced our own evaluating
function Cumulative Effectiveness and a derivative Utilization Factor.
Finally, we implemented an experimentation system to test these
algorithms on different sets of tasks to allocate. In this article
there is a short analysis of series of experiments conducted on three
different classes of task sets: small tasks, mixed tasks and large
tasks.

#e1#

#s1#
#abstract1409.txt#


DAG scheduling is of great importance to optimal distribution of tasks
in parallel and distributed systems. In this paper a novel approach to
DAG scheduling, utilizing learning automata across distributed systems,
is proposed. The learning process begins with an initial population of
randomly generated learning automata. Each automaton by itself
represents a stochastic scheduling. The scheduling is optimized within
a learning process. Compared with current genetic approaches to DAG
scheduling better results are achieved. The main reason underlying this
achievement is that an evolutionary approach such as genetics looks for
the best chromosomes within genetic populations whilst in the approach
presented in this paper learning automata is applied to find the most
suitable position for the genes in addition to looking for the best
chromosomes. The scheduling resulted from applying our scheduling
algorithm to some benchmark task graphs are compared with the existing
ones.

#e1#

#s1#
#abstract1410.txt#


In this report, we experimentally demonstrate that single platinum
nanoparticles exhibit the necessary catalytic activity for the
optically induced reduction of H[AuCl/sub 4/] complexes to elemental
gold. This finding is exploited for the parallel Au encapsulation of
FePt nanoparticles arranged in a self-assembled two-dimensional array.
Magnetic force microscopy reveals that the thin gold layer formed on
the FePt particles leads to a strongly increased long-term stability of
their magnetization under ambient conditions.

#e1#

#s1#
#abstract1411.txt#


The purpose of content-based image retrieval (CBIR) is to retrieve,
from real data stored in a database, information that is relevant to a
query. In remote sensing applications, the wealth of spectral
information provided by latest-generation (hyperspectral) instruments
has quickly introduced the need for parallel CBIR systems able to
effectively retrieve features of interest from ever-growing data
archives. To address this need, this paper develops a new parallel CBIR
system that has been specifically designed to be run on heterogeneous
networks of computers (HNOCs). These platforms have soon become a
standard computing architecture in remote sensing missions due to the
distributed nature of data repositories. The proposed heterogeneous
system first extracts an image feature vector able to characterize
image content with sub-pixel precision using spectral mixture analysis
concepts, and then uses the obtained feature as a search reference. The
system is validated using a complex hyperspectral image database, and
implemented on several networks of workstations and a Beowulf cluster
at NASA's Goddard Space Flight Center. Our experimental results
indicate that the proposed parallel system can efficiently retrieve
hyperspectral images from complex image databases by efficiently
adapting to the underlying parallel platform on which it is run,
regardless of the heterogeneity in the compute nodes and communication
links that form such parallel platform. Copyright c 2009 John Wiley &
Sons, Ltd.

#e1#

#s1#
#abstract1415.txt#


Catalyst- and mask-free grown GaN nanorods have been investigated using
transmission electron microscopy (TEM), scanning transmission electron
microscopy (STEM) and energy filtered transmission electron microscopy
(EFTEM). The nanorods were grown on nitridated r-plane sapphire
substrates in a molecular beam epitaxy reactor. We investigated samples
directly after the nitridation and after the overgrowth of the
structure with GaN. High resolution transmission electron microscopy
(HRTEM) and EFTEM revealed that AlN islands have formed due to
nitridation. After overgrowth, the AlN islands could not be observed
any more, neither by EFTEM nor by Z-contrast imaging. Instead, a smooth
layer consisting of AlGaN was found. The investigation of the overgrown
sample revealed that an a-plane GaN layer and GaN nanorods on top of
the a-plane GaN have formed. The nanorods reduced from top of the
a-plane GaN towards the a-plane GaN/sapphire interface suggesting that
the nanorods originate at the AlN islands found after nitridation.
However, this could not be shown unambiguously. The number of threading
dislocations in the nanorods was very low. The analysis of the
epitaxial relationship to the a-plane GaN showed that the nanorods grew
along the [000-1] direction, and the [1-100] direction of the rods was
parallel to the [0001] direction of the a-plane GaN.

#e1#

#s1#
#abstract1417.txt#


Optical parallel processing enables electronic dispersion compensation
in coherent receivers at symbol rates above the employed electronics'
speed limit. Up to 64 Gbaud QPSK (128 Gb/s) are processed and
transmitted up to 610 km-SSMF without optical dispersion compensation.

#e1#

#s1#
#abstract1418.txt#


Parallel-to-cascading optical code label processing is proposed and
experimentally demonstrated in a 16-node setup. The physical structure
for multi-bit label's processing is significantly simplified, since
only one encoder, one decoder and one photodiode are needed.

#e1#

#s1#
#abstract1419.txt#


The Discrete Fourier Transform (DFT) is a mathematical procedure that
stands at the center of the processing that takes place inside a
Digital Signal Processor. It has been known and argued through the
literatures that the Fast Fourier Transform (FFT) is useless in
detecting a specific frequency in a monitored signal because most of
the computed results are ignored. In this paper we will present an
efficient FFT based method to detect specific frequencies in a
monitored signal which is compared to the most frequently used method
"the Goertzel's Algorithm". Parallel implementation structure show a
fast computation method compared to the Goertzel's algorithm.
Computational speedup gains of r using radix-r butterfly are shown.

#e1#

#s1#
#abstract1420.txt#


Phase-based direction-of-arrival (DOA) techniques estimate the phase
difference between the output signal of a conventional beamformer and
that of a related parallel beamformer. Since the two beamformers share
the same weights in a specific known manner, the DOA estimators are
affected by any phase mismatch of the beamformer weights relative to
the direction vector of the signal of interest. This paper investigates
the sensitivity of two phase-based DOA estimators to such beamformer
mismatch via a theoretical analysis, and presents an example computer
simulation to illustrate their properties.

#e1#

#s1#
#abstract1421.txt#


Various information systems are widely used in information society era,
and the demand for highly dependable system is increasing year after
year. However, software testing for such a system becomes more
difficult due to the enlargement and the complexity of the system. In
particular, it is too difficult to test parallel and distributed
systems sufficiently although dependable systems such as
high-availability servers usually form parallel and distributed
systems. To solve these problems, we proposed a software testing
environment for dependable parallel and distributed system using the
cloud computing technology, named D-Cloud. D-Cloud includes Eucalyptus
as the cloud management software, and FaultVM based on QEMU as the
virtualization software, and D-Cloud frontend for interpreting test
scenario. D-Cloud enables not only to automate the system configuration
and the test procedure but also to perform a number of test cases
simultaneously, and to emulate hardware faults flexibly. In this paper,
we present the concept and design of D-Cloud, and describe how to
specify the system configuration and the test scenario. Furthermore,
the preliminary test example as the software testing using D-Cloud was
presented. Its result shows that D-Cloud allows to set up the
environment easily, and to test the software testing for the
distributed system.

#e1#

#s1#
#abstract1422.txt#


The York Extensible Testing Infrastructure (YETI) is an automated
random testing tool that allows to test programs written in various
programming languages. While YETI is one of the fastest random testing
tools with over a million method calls per minute on fast code, testing
large programs or slow code -- such as libraries using intensively the
memory -- might benefit from parallel executions of testing sessions.
This paper presents the cloud-enabled version of YETI. It relies on the
Hadoop package and its map/reduce implementation to distribute tasks
over potentially many computers. This would allow to distribute the
cloud version of YETI over Amazon's Elastic Compute Cloud (EC2).

#e1#

#s1#
#abstract1423.txt#


Many embedded and distributed applications are based on processing
nodes that perform parallel processing tasks. Unfortunately, it is
difficult to evaluate the overall behaviour of this kind of
applications because the overall behaviour consists of 1) the
execution-paths of asynchronous processing nodes and of 2) messages
that either activate or deactivate processing nodes to perform parallel
processing tasks. In order to facilitate behaviour and reliability
evaluation of applications doing parallel processing, we developed a
method that: 1) is capable of composing an overall representation for
parallel behaviours and recognizing both the defined use cases and
undetermined behaviours from this representation and 2) supports
calculation of use case-specific reliability values for components. In
this paper, we describe the method, present a ComponentBee tool that
implements the method and supports behaviour and reliability evaluation
of multithreaded Java applications, and finally demonstrate the use of
the method with a case study.

#e1#

#s1#
#abstract1424.txt#


This paper presents clustering of fatigue features resulted from the
segmentation of the SAESUS time series data. The segmentation process
was based on the Morlet wavelet coefficient amplitude level which
produces 49 segments that each has an overall fatigue damage.
Observation of the fatigue damage and the wavelet coefficients was made
on each segment. In the end of the process, the segments were clustered
into three clusters in order to identify any improvements in the data
scattering for fatigue data clustering prospects. This algorithm
produced a more reliable and suitable method of segment by segment
analysis for fatigue strain signal segmentation. According to the
findings, the higher Morlet wavelet coefficient presented damaging
segment, otherwise, it was undamaging segment. It indicated that the
relationship between the Morlet wavelet coefficient and the fatigue
damage was strong and parallel.

#e1#

#s1#
#abstract1425.txt#


Task scheduling is an essential aspect of parallel processing system.
This problem assumes fully connected processors and ignores contention
on the communication links. However, as arbitrary processor network
(APN), communication contention has a strong influence on the execution
time of a parallel application. In this paper, we propose
multi-objective genetic algorithm to solve task scheduling problem with
time constraints in unstructured heterogeneous processors to find the
scheduling with minimum makespan and total tardiness. To optimize
objectives, we use Pareto front based technique, vector based method.
In this problem, just like tasks, we schedule messages on suitable
links during the minimization of the makespan and total tardiness. To
find a path for transferring a message between processors we use
classic routing algorithm. We compare our method with BSA method that
is a well known algorithm. Experimental results show our method is
better than BSA and yield better makespan and total tardiness.

#e1#

#s1#
#abstract1427.txt#


A pushbroom MSI sensor collects image data from the ground, parallel to
the flight path, at a specific point angle. Images taken at two
instances of time while the scanner moves with the platform usually
have a spatial offset, and need to be registered before they can be
compared for any changes between the two images. Moving target
detection is a special case of change detection that requires the time
between frames be small enough that a moving vehicle remains in close
proximity on the two frames. We propose an algorithm for the detection
of moving targets in a multi-band line scanning pushbroom sensor.
Ideally, change detection works best when images have the same spectral
bandwidth and are perfectly registered to one another, since
differencing the two images automatically removes most of the common
background signal. However, this is not always the case. For example,
the sensor considered here has different bandwidths for its component
bands, and since it is a line-scanner it is much more challenging
regarding image registration than a framed-based scanner. In this
study, we will use simulated data of the same bandwidth to demonstrate
the fundamental algorithm of detection and velocity calculation. The
velocity calculation is that of distance divided by time; but,
depending on the focal plane layout and other operating considerations
and conversion between image space to physical units, this calculation
is not as simple as it seems. We will also discuss our effort in
applying our algorithm to real line-scan imagery, of different
bandwidths in the two channels. We will show the extra image processing
efforts needed to make it work, and show some of the test results.

#e1#

#s1#
#abstract1431.txt#


In this paper, we propose a novel face recognition method based on
local appearance feature extraction using dual-tree complex wavelet
transform (DT-CWT). It provides a local multiscale description of
images with good directional selectivity and invariance to shifts and
in-plane rotations. In the dual-tree implementation, two parallel
discrete wavelet transform (DWT) with different lowpass and highpass
filters in different scales were used. The linear combination of
subbands generated by two parallel DWT is used to generate 6 different
directional subbands with complex coefficients. It is insensitive to
illumination variations and facial expression changes. 2-D dual-tree
complex wavelet transform is less redundant and computationally
efficient. The local DT-CWT coefficients are used to extract the facial
features which improve the face recognition with small sample size in
less computation. The local features based methods have been
successfully applied to face recognition and achieved state-of-the-art
performance. Normally most of the local appearance based methods the
facial features are extracted from several local regions and
concatenated into an enhanced feature vector as a face descriptor. In
this approach we divide the face into several (m * m) non-overlapped
parallelogram blocks instead of square or rectangle blocks. The local
mean, standard deviation and energy of complex wavelet coefficients are
used to describe the face image. Experiments, on two well-known
databases, namely, Yale and ORL databases, shows the Local DT-CWT
approach performs well on illumination, expression and perspective
variant faces with single sample compared to PCA and global DT-CWT.
Furthermore, in addition to the consistent and promising classification
performances, our proposed Local DT-CWT based method has a really low
computational complexity.

#e1#

#s1#
#abstract1432.txt#


Multiple sequence alignment is the most common task in computational
biology. This multiple sequence alignment is computationally difficult
and classified as a NP-Hard problem; so approximate algorithm(s) are
generally required for most multiple alignment tasks. The Molecular
Biologist may require the alignment of thousands of sequences that each
can be of many hundreds of amino acids or even several millions of
nucleotides. The approximation algorithm requires a long processing
period of time to compute near optimal alignment. Thus, one step to
reduce the processing time is to parallelize the algorithm. In order to
have solution over parallelism method, we can either use expensive
multiprocessor programming or cheaper cluster/Grid programming.
Multiprocessor systems are specialized expensive hardware and are not
commonly available. An alternative cheapest way is to use either a
computer cluster or a Computing Grid. A cluster can be used for amino
acid sequences and will be very slow for multiple sequence alignment of
DNA molecule. So, the computing grid is the only cheapest alternative
for performing multiple sequence alignment of DNA molecules. We have
designed an efficient grid scheduler to perform the parallel tasks in
grid that minimizes the communication cost and time complexity and also
implemented parallel algorithm on computing grid. The experimental
results show enhanced speedup.

#e1#

#s1#
#abstract1433.txt#


An automated general purpose method is introduced for computing a
rigorous estimate of a bounded region in R/sup n/ whose points satisfy
a given property. The method is based on calculations conducted in
interval arithmetic and the constructed approximation is built of
rectangular boxes of variable sizes. An efficient strategy is proposed,
which makes use of parallel computations on multiple machines and
refines the estimate gradually. It is proved that under certain
assumptions the result of computations converges to the exact result as
the precision of calculations increases. The time complexity of the
algorithm is analyzed, and the effectiveness of this approach is
illustrated by constructing a lower bound on the set of parameters for
which an over-compensatory nonlinear Leslie population model exhibits
more than one attractor, which is of interest from the biological point
of view. This paper is accompanied by efficient and flexible software
written in C++ whose source code is freely available at
http://www.pawelpilarczyk.com/parallel/.

#e1#

#s1#
#abstract1435.txt#


While OpenFlow has potential to open control of the network, only one
researcher can innovate on the network at a time. What is required is a
way to divide, or slice, network resources so that researchers and
network administrators can use them in parallel. Network slicing
implies that actions in one slice do not negatively affect other
slices, even if they share the same underlying physical hardware.
FlowVisor, a special purpose OpenFlow controller that allows multiple
researchers to run experiments safely and independently on the same
production OpenFlow network is introduced in this paper.
Architecturally, FlowVisor acts as a transparent virtualization layer
between OpenFlow switches and controllers. Network devices generate
OpenFlow protocol messages, which go to the FlowVisor and are then
routed by network slice to the appropriate researcher.

#e1#

#s1#
#abstract1438.txt#


This paper studies the parallel machines bi-criteria scheduling problem
(PMBSP) in a deteriorating system. Sequencing and scheduling problems
(SSP) have seldom considered the two phenomena concurrently. This paper
discusses the parallel machines scheduling problem with the effects of
machine and job deterioration. By the machine deterioration effect, we
mean that each machine deteriorates at a different rate. This
deterioration is considered in terms of cost which depends on the
production rate, the machine's operating characteristics and the kind
of work done by each machine. Moreover, job processing times are
increasing functions of their starting times and follow a simple linear
deterioration. The objective functions are minimizing total tardiness
and machine deteriorating cost. The problem of total tardiness on
identical parallel machines is NP-hard, thus the problem with machine
deteriorating cost as an additional term is also NP-hard. We propose
the LP-metric method to show the importance of our proposed
multi-objective problem. A metaheuristic algorithm is developed to
locate optimal or near optimal solutions based on a Tabu search
mechanism. Numerical examples are presented to show the efficiency of
this model. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1440.txt#


Real-time video processing is a rapidly evolving field with growing
applications in science and engineering. Portable video processing
systems require design which reduces the power, memory usage, and
resource utilization while maintaining real-time operation. General
system architecture for real-time video filtering and overlay character
generation, based on a Field Programmable Gate Array (FPGA) is
presented and evaluated. After initial configuration over the parallel
port connection, the video codec dumps a stream of pixels in ITU-R
BT.656 format, 8-bits per component with sync signals, using a sampling
clock of 27 MHz. This stream is then fed to a series of filters, format
conversion and overlay character generator. Finally, the Video Graphics
Array (VGA) generator receives the pixel stream and displays it on a
monitor. The implemented design includes filters of different functions
and character generator with keyboard interface. The real-time of
application like stock ticker and time display are also demonstrated.
The VHDL codes for the architecture were synthesized using Xilinx ISE
10.1 and targeted for Xilinx Spartan3E FPGA. The results show that this
design gives good performance with short processing time, low resource
utilization, small power consumption and memory usage.

#e1#

#s1#
#abstract1442.txt#


In parallel with studies, a lot of extra activities need to be fitted
in a student's schedule. Frequently, excessive workload results in poor
performance or in failing to finish the studies. The problem is more
severe in lifelong learning, where students are professionals with
family duties. So, the need of making informative decisions as of
whether taking a specific course fits into a student's schedule is of
great importance. This paper illustrates a system, called EDUPLAN and
being currently under development, which aims at helping the student to
make intelligent management of her time. EDUPLAN aims at informing the
student as for which learning objects can fit her schedule or not, as
well as at organizing her time. This can be achieved using scheduling
algorithms and a description of the user's tasks and events. In the
paper we also extend the LOM 1484.12.3 trade -2005 ontology with
classes that can be used to describe the temporal distribution of the
workload of any learning object. Finally, we provide EDUPLAN's
architecture, being built around the existing SELFPLANNER intelligent
calendar application.

#e1#

#s1#
#abstract1443.txt#


This paper presents an advanced 640 * 480 (VGA) IRFPA based on uncooled
microbolometers with a pixel-pitch of 25 mu m developed by
Fraunhofer-IMS. The IRFPA is designed for thermal imaging applications
in the LWIR (8 .. 14 mu m) range with a full-frame frequency of 30 Hz
and a high sensitivity with NETD # 100 mK @ f/1. A novel readout
architecture which utilizes massively parallel on-chip Sigma-Delta-ADCs
located under the microbolometer array results in a high performance
digital readout. Sigma-Delta-ADCs are inherently linear. A high
resolution of 16 bit for a secondorder Sigma-Delta-modulator followed
by a third-order digital sinc-filter can be obtained. In addition to
several thousand Sigma-Delta-ADCs the readout circuit consists of a
configurable sequencer for controlling the readout clocking signals and
a temperature sensor for measuring the temperature of the IRFPA. Since
packaging is a significant part of IRFPA's price Fraunhofer-IMS uses a
chip-scaled package consisting of an IR-transparent window with
antireflection coating and a soldering frame for maintaining the
vacuum. The IRFPAs are completely fabricated at Fraunhofer-IMS on 8"
CMOS wafers with an additional surface micromachining process. In this
paper the architecture of the readout electronics, the packaging, and
the electro-optical performance characterization are presented.

#e1#

#s1#
#abstract1444.txt#


Besides resolution, an important performance parameter of a FIR camera
is the sensitivity. It depends on the sensitivity of the detector array
itself and the characteristics of the optic. The effects of the optic
are considerably driven by the f-number, with high values resulting in
decreased sensitivity, but providing the possibility for simple lens
design and cheaper production costs. In this contribution 4 different
sensor setups with different optics are evaluated for their impact on
the performance of trained pedestrian classifiers. To overcome the
expensive and time consuming process of ground truth generation for
multiple sensors, an approach for reusing available high sensitivity
reference data is presented. Classifiers are trained on specially
transformed reference data with characteristics of sensors with
degraded sensitivity. For the evaluation of the classifiers, data of
real world road scenarios is collected simultaneously with the target
sensors mounted in parallel in a test vehicle, following a detailed
script for recording a pedestrian scene test catalogue. This allows for
a direct analysis and comparison of the different sensors and their
impact on the detection performance.

#e1#

#s1#
#abstract1445.txt#


An optical SAR processor prototype exhibiting real-time and fine
sampling capabilities has been successfully developed and tested.
Synthetic Aperture Radar (SAR) images are typically processed digitally
applying dedicated Fast Fourier Transform (FFT) algorithms. These
operations are time consuming and require a large amount of processing
power and are often performed in one dimension at a time. A true two
dimensional Fourier transform may be instead performed through optics,
as optical processing provides inherent parallel computing
capabilities. By processing the azimuth and slant range directions
simultaneously, a reduction in processing time and power is achieved.
In addition, the configuration of the optics is such that high
resolution images may be obtained at no additional processing cost. The
optical SAR processor is also designed to adapt to SAR system parameter
changes. It has the capability to produce full Envisat/ASAR scenes from
the various image mode swaths (IS1-IS7) within tens of seconds. This
paper reviews the design of the real-time high resolution optical SAR
processor prototype and discusses the results of images reconstructed
from simulated point targets as well as from Envisat/ASAR data sets.

#e1#

#s1#
#abstract1446.txt#


With sensor technologies rapidly improving, the need to process
increasingly larger data sets is becoming the main bottleneck in many
real time applications associated with persistence surveillance such as
VideoSAR and volumetric SAR imaging. In many instances, the image
fidelity is of utmost importance which can have implications when
choosing the appropriate algorithm to generate the desired data
products. The performance improvements afforded by algorithms such as
the fast back projection (FBP) algorithm prove attractive for such
environments. Unfortunately, even though the FBP algorithm is
magnitudes faster than a traditional back projection algorithm it is
still incapable of meeting the strict requirements of some of the
aforementioned real time applications. However, the emergence of
general purpose graphical processing units (GPGPUs) in recent years
have afforded many scientific fields orders of magnitudes improvement
in performance for a large variety of applications. This is also the
case for the FBP algorithm. By distributing the processing across 480
processing cores located on a single video card, it possible to achieve
substantial performance improvements compared to the serial FBP
algorithm. Considering that many PCs are capable of housing three to
four video cards, it is possible to obtain more than two orders of
magnitude improvement in performance with the parallel approach. This
technology provides the ability to process enormous datasets in the
field without the need of supercomputers that have to date been the
only means of keeping pace with the incoming data.

#e1#

#s1#
#abstract1447.txt#


Modern image enhancement techniques have been shown to be effective in
improving the quality of imagery. However, the computational
requirements of applying such algorithms to streams of video in
real-time often cannot be satisfied by standard microprocessor-based
systems. While a scaled solution involving clusters of microprocessors
may provide the necessary arithmetic capacity, deployment is limited to
data-center scenarios. What is needed is a way to perform these
techniques in real time on embedded platforms. A new paradigm of
computing utilizing special-purpose commodity hardware including
Field-Programmable Gate Arrays (FPGAs) and Graphics Processing Units
(GPU) has recently emerged as an alternative to parallel computing
using clusters of traditional CPUs. Recent research has shown that for
many applications, such as image processing techniques requiring
intense computations and large memory spaces, these hardware platforms
significantly outperform microprocessors. Furthermore, while
microprocessor technology has begun to stagnate, GPUs and FPGAs have
continued to improve exponentially. FPGAs, flexible and powerful, are
best targeted at embedded, low-power systems and specific applications.
GPUs, cheap and readily available, are available to most users through
their standard desktop machines. Additionally, as fabrication scale
continues to shrink, heat and power consumption issues typically
limiting GPU deployment to high-end desktop workstations are becoming
less of a factor. The ability to include these devices in embedded
environments opens up entire new application domains. In this paper, we
investigate two state-of-the-art image processing techniques,
super-resolution and the average-bispectrum speckle method, and compare
FPGA and GPU implementations in terms of performance, development
effort, cost, deployment options, and platform flexibility.

#e1#

#s1#
#abstract1448.txt#


In this paper we present a drift-correcting template update strategy
for precisely tracking a feature point in 2D image sequences. The
proposed strategy greatly complements one of the latest published
template update strategies by incorporating a robust non-rigid image
registration step. Previous strategies use the first template to
correct drifts in the current template; however, the drift still builds
up when the first template becomes different from the current one
particularly in a long image sequence. In our strategy the first
template is updated timely when it is revealed to be quite different
from the current template and henceforth the updated first template is
used to correct template drifts in subsequent frames. Our method runs
fast on a 3.0 GHz desktop PC, using about 0.03 s on average to track a
feature point in a frame (under the assumption of a general affine
transformation model, 61 * 61 pixels in template size) and less than
0.1 s to update the first template. The proposed template update
strategy can be implemented either serially or in parallel.
Quantitative evaluation results show the proposed method in precision
tracking of a distinctive feature point whose appearance is constantly
changing. Qualitative evaluation results show that the proposed method
has a more sustained ability to track a feature point than two previous
template update strategies. We also revealed the limitations of the
proposed template update strategy by tracking feature points on a
human's face. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1449.txt#


Background: In recent years, the demand for computational power in
computational biology has increased due to rapidly growing data sets
from microarray and other high-throughput technologies. This demand is
likely to increase. Standard algorithms for analyzing data, such as
cluster algorithms, need to be parallelized for fast processing.
Unfortunately, most approaches for parallelizing algorithms largely
rely on network communication protocols connecting and requiring
multiple computers. One answer to this problem is to utilize the
intrinsic capabilities in current multi-core hardware to distribute the
tasks among the different cores of one computer. Results: We introduce
a multi-core parallelization of the k-means and k-modes cluster
algorithms based on the design principles of transactional memory for
clustering gene expression microarray type data and categorial

#e1#

#s1#
#abstract1451.txt#


A rapid nondestructive measurement method for determining the total
viable count of chilled pork was studied. Chilled pork samples were
purchased from supermarket and then stored in refrigerator at 4 degrees
C. Every 24 hours, hyperspectral images were collected from the chilled
pork samples in 400-1100nm region, in parallel total viable counts were
obtained by classical microbiological plating methods. The 3-parameter
modified lorentzian distribution function was applied to fit the
scattering profiles of all samples and the fitting results were
satisfactorily high in region 470-943 nm. Then the parameters extracted
were used to establish PLSR models. The prediction results for the
parameter a, b, c, b*c are 0.945, 0.918, 0.919, 0.935 respectively. The
study show that the hyperspectral technology can accurately tracks the
increase of total viable count of chilled pork during 2-14 days storage
at 4 degrees C, and so indicate it a valid tool for assessing the
quality and safety properties of chilled pork rapidly and
nondestructively in the future.

#e1#

#s1#
#abstract1454.txt#


Various methods for fractional delay in the frequency domain are
proposed and discussed as special cases of interpolation applied to
audio signal beamforming applications. The techniques for time shifting
in the frequency domain are well known, particularly with the discrete
Fourier transform. However, the application details for windowed
signals, especially for the purpose of time compensation in beamforming
applications, have seldom been studied and discussed in the literature.
Nowadays, with the growing interest in array processing, fractional
delay may be widely applied. The methods discussed are especially
suited for frequency processing procedures that consist of windowing +
transform with the usual frequency transforms-the discrete Fourier
transform (DFT) and the discrete cosine transform (DCT). In parallel,
the effects on beamforming precision caused by time-compensation errors
when fractional delay is not applied are analyzed for the case of
uniform linear microphone arrays.

#e1#

#s1#
#abstract1455.txt#


In this article a systematic approach of modelling and control for a
parallel robotic manipulator is presented. Regarding the framework of
structured analysis of dynamical systems the derivation of a
differential-algebraic model of the mechanical system is
straightforward. Using some differential-geometric considerations based
on invariant manifolds and the definition of fictitious additional
input and output variables a suitable state feedback can be constructed
which transforms the differential-algebraic representation into a
state-space model for the robotic manipulator. On this basis a
classical two-degree-of-freedom (2-DOF) control structure has been
designed using the well-known input-output linearization and a linear
time-variant Kalman filter-based output feedback. Finally, the control
structure including a friction compensation is applied to the robotic
system in the laboratory which shows the practical applicability of the
proposed procedure.

#e1#

#s1#
#abstract1456.txt#


Photomask degradation via haze defect formation is an increasing
troublesome yield problem in the semiconductor fab. Wafer inspection is
often utilized to detect haze defects due to the fact that it can be a
bi-product of process control wafer inspection; furthermore, the
detection of the haze on the wafer is effectively enhanced due to the
multitude of distinct fields being scanned. In this paper, we
demonstrate a novel application for enhancing the wafer inspection
tool's sensitivity to haze defects even further. In particular, we
present results of bright field wafer inspection using the on several
photo layers suffering from haze defects. One way in which the enhanced
sensitivity can be achieved in inspection tools is by using a double
scan of the wafer: one regular scan with the normal recipe and another
high sensitivity scan from which only the repeater defects are
extracted (the non-repeater defects consist largely of noise which is
difficult to filter). Our solution essentially combines the double scan
into a single high sensitivity scan whose processing is carried out
along two parallel routes (see Fig. 1). Along one route, potential
defects follow the standard recipe thresholds to produce a defect map
at the nominal sensitivity. Along the alternate route, potential
defects are used to extract only field repeater defects which are
identified using an optimal repeater algorithm that eliminates "false
repeaters". At the end of the scan, the two defect maps are merged into
one with optical scan images available for all the merged defects. It
is important to note, that there is no throughput hit; in addition, the
repeater sensitivity is increased relative to a double scan, due to a
novel runtime algorithm implementation whose memory requirements are
minimized, thus enabling to search a much larger number of potential
defects for repeaters. We evaluated the new application on photo wafers
which consisted of both random and haze defects. The evaluation
procedure involved scanning with three different recipe types: Standard
Inspection: Nominal recipe with a low false alarm rate was used to scan
the wafer and repeaters were extracted from the final defect map. Haze
Monitoring Application: Recipe sensitivity was enhanced and run on a
single field column from which on repeating defects were extracted.
Enhanced Repeater Extractor: Defect processing included the two
parallel routes: a nominal recipe for the random defects and the new
high sensitive repeater extractor algorithm. The results showed that
the new application (recipe #3) had the highest capture rate on haze
defects and detected new repeater defects not found in the first two
recipes. In addition, the recipe was much simpler to setup since
repeaters are filtered separately from random defects. We expect that
in the future, with the advent of mask-less lithography and EUV
lithography, the monitoring of field and die repeating defects on the
wafer will become a necessity for process control in the semiconductor
fab.

#e1#

#s1#
#abstract1457.txt#


This paper has presented some engineering work on developing a
micropipeline blocksorter. The work presented in this paper
demonstrates that VHDL can be used to describe the behaviour of
micropipelined systems. It also shows a comparison of 2-phase and
4-phase implementations in transistor count, speed, and energy. Though
the nature of the work is mainly engineering, there are some
significant new insights gained in the course of the work.

#e1#

#s1#
#abstract1458.txt#


The WK-recursive network proposed by Vecchia and Sanges [7] is widely
used in the design and implementation of local area networks and
parallel processing architectures. It provides a high degree of
regularity and scalability, which very well conforms to a modular
design and realization of distributed systems involving large number of
computing elements. In this paper, we investigate the routing of a
message on the WK-recursive network, that is key to the performance of
this network. We present an efficient shortest path algorithm on the
WK-recursive network, which is better that Chen and Duh [2] in terms of
designing complexity.

#e1#

#s1#
#abstract1459.txt#


This paper reports an efficient implementation of the Discrete Wavelet
Packet Transform with hardware acceleration. This design relies on the
implementation of the word-serial pipeline architecture and filter
parallelism to minimize the processing time and maximize performance.
The acceleration, working two times faster than Architecture reported
in is achieved by designing parallel high-pass and low-pass filters at
each tree level using internal FPGA multipliers. The architecture can
be implemented for any filter with different orders. The performance
evaluation is made with the AT/sup 2/ figure of merit. The results show
that the implementation reported provides an AT/sup 2/ figure that is
consistently smaller than 0.5 outperforming previous implementations
reported in. This high speed architecture can be used to implement the
Direct Wavelet Packet Transform at any tree level and is suitable for
real-time applications.

#e1#

#s1#
#abstract1460.txt#


In this paper, based on the word-serial pipeline architecture and
parallel filter processing, a new architecture for direct and inverse
wavelet packet transforms is introduced. This architecture increase the
speed of the wavelet packet transforms. In this design a word-serial
architecture able to compute a complete wavelet packet transform (WPT)
binary tree in an on-line fashion, but easily configurable in order to
compute any required WPT sub tree, is proposed. In this architecture, a
high-pass filter and a low-pass filter are used concurrently, in order
to compute the new coefficients. This architecture is suitable for the
high-speed on-line applications. With this architecture, the speed of
the wavelet packet transforms is increased with a factor two, but the
occupied area of the circuit is less than double. This architecture can
be applied to any levels tree structure with any filter coefficients
length.

#e1#

#s1#
#abstract1461.txt#


Data Warehouses are databases used in Business Intelligence systems as
a data source to develop analytical applications. These applications
consist of multidimensional analyses of data and allow decisional
makers to improve the business processes of the Information System.
Since multidimensional analyses require to aggregate data on several
attributes, techniques based on approximate query answering have been
introduced in order to reduce the response time. These techniques use,
as a data source, a synopsis of the data stored in the Data Warehouse.
In this paper, a parallel algorithm for the computation of data
synopsis is presented.

#e1#

#s1#
#abstract1462.txt#


The programmable graphics processing unit (GPU) is employed to
accelerate the unconditionally stable Crank-Nicolson finite-difference
time-domain (CN-FDTD) method for the analysis of microwave circuits. In
order to efficiently solve the linear system from the CN-FDTD method at
each time step, both the sparse matrix vector product (SMVP) and the
arithmetic operations on vectors in the bi-conjugate gradient
stabilized (Bi-CGSTAB) algorithm are performed with multiple processors
of the GPU. Therefore, the GPU based BI-CGSTAB algorithm can
significantly speed up the CN-FDTD simulation due to parallel computing
capability of modern GPUs. Numerical results demonstrate that this
method is very effective and a speedup factor of 10 can be achieved.

#e1#

#s1#
#abstract1463.txt#


The packet routing problem, i.e., the problem to send a given set of
unit-size packets through a network on time, belongs to one of the most
fundamental routing problems with important practical applications,
e.g., in traffic routing, parallel computing, and the design of
communication protocols. The problem involves critical routing and
scheduling decisions. One has to determine a suitable (short)
origin-destination path for each packet and resolve occurring conflicts
between packets whose paths have an edge in common. The overall aim is
to find a path for each packet and a routing schedule with minimum
makespan. A significant topology for practical applications are grid
graphs. In this paper, we therefore investigate the packet routing
problem under the restriction that the underlying graph is a grid. We
establish approximation algorithms and complexity results for the
general problem on grids, and under various constraints on the start
and destination vertices or on the paths of the packets.

#e1#

#s1#
#abstract1467.txt#


MapReduce model is a new parallel programming model initially developed
for large-scale web content processing. Data analysis meets the issue
of how to do calculation over extremely large dataset. The arrival of
MapReduce provides a chance to utilize commodity hardware for massively
parallel data analysis applications. The translation and optimization
from relational algebra operators to MapReduce programs is still an
open and dynamic research field. In this paper, we focus on a special
type of data analysis query, namely, multiple group by query. We first
study the communication cost of MapReduce model, then we give an
initial implementation of multiple group by query. We then propose an
optimized version which addresses and improves the communication cost
issues. Our optimized version shows a better accelerating ability and a
better scalability than the other version.

#e1#

#s1#
#abstract1468.txt#


Scheduling workflow applications in grid environments is a great
challenge, because it is an NP-complete problem. Many heuristic methods
have been presented in the literature and most of them deal with a
single workflow application at a time. In recent years, there are
several heuristic methods proposed to deal with concurrent workflows or
online workflows, but they do not work with workflows composed of
data-parallel tasks. In this paper, we present an online scheduling
approach for multiple mixed-parallel workflows in grid environments.
The proposed approach was evaluated with a series of simulation
experiments and the results show that the proposed approach delivers
good performance and outperforms other methods under various workloads.

#e1#

#s1#
#abstract1472.txt#


In order to investigate the double-pulse ablation mechanism, two
parallel but non-collinear laser beams, delayed with respect to each
other by 1 mu s, were focussed on an aluminium sample, so that a
lateral distance of 600 microns exists between the centres of the two
craters and no superposition of the laser-ablation zones is present.
The use of such configuration results in a signal and in a plasma mass
enhancement with respect to the single-pulse case almost equal to that
obtained in the double-pulse collinear case. However, such a
non-collinear geometry evidences a much more effective drilling of the
surface. Such unexpected drilling seems to be related to a hydrodynamic
drainage out of aerosol and molten material, hindering its
re-deposition in and around the crater.

#e1#

#s1#
#abstract1473.txt#


We have developed an algorithm, a set of parallel programs for seismic
wave field simulation and have carried out a number of test
computations on clusters of the Siberian Supercomputer Center (SB RAS),
in order to choose the optimal parallel scheme.Mathematical modeling of
elastic wave propagation from a point source in the 3D-models of
elastic media, characteristic of mud volcanoes, was performed. Some
results of processing the data from the vibroseismic experiments on the
Karabetov Mountain mud volcano are discussed. The results of numerical
and natural experiments are compared. The paper was prepared on the
basis of the authors' report at the International Conference on
Parallel Computing Technologies (PaVT-2010; http://agora.guru.ru/pavt).

#e1#

#s1#
#abstract1474.txt#


The paper deals with the technology of generation and identification of
2D IIR-filter characteristics. For the characteristic identification,
small test fragments of an image are used; these fragments are formed
from a distorted image using the prior information on the geometric
form of objects. 2D IIR-filter is generated as parallel casual filters
with a mask in the form of a quadrant. A distributed system
architecture is proposed with consideration of data parallelism for
high-resolution images and the parallel message passing with regard to
the IIR-filter structure. The paper was prepared on the basis of the
authors' report at the International Conference on Parallel Computing
Technologies (PaVT-2010; http://agora.guru.ru/pavt).

#e1#

#s1#
#abstract1475.txt#


Most traditional methods for /i T//sub 1/ map estimation in MRI with
fast low-angle-shot sequences are aimed at high efficiency by
compromising the fitting accuracy. In this paper, the /i fundamental/
problem of parameter estimation in fast low-angle-shot MRI was
re-examined, and an accurate and fast optimization approach, named
concatenated optimization for parameter estimation, was proposed for
the regression of data points acquired with multiple flip angles. The
initial estimation of /i T//sub 1/ was obtained from the linear
regression, followed by the constrained nonlinear regression based on
the initial estimates. This heterogeneous initialization strategy
improves the fitting accuracy and reduces the computational time. A
computationally efficient implementation of concatenated optimization
for parameter estimation was achieved based on the graphic processing
unit, named as concatenated optimization for parameter estimation
graphic processing unit. In experimental comparison with Fram's method
and the Fitter Tool in Jim, the proposed methods are capable of
achieving significantly higher efficiency and more accurate
estimations. Magn Reson Med 63:1431-1436, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract1476.txt#


Hypointense band artifacts occur at intersections of nonparallel
imaging planes in rapidly acquired MR images; quantitative or numerical
analysis of these bands and strategies to mitigate their appearance
have largely gone unexplored. The magnetization evolution in the
different regions of multiplanar images was simulated for three common
rapid steady-state techniques (spoiled gradient echo, steady state free
precession, balanced steady state free precession). Saturation banding
was found to be highly dependent on the pulse sequence, acquisition
time, and phase-encoding order. Encoding the center of /i k/-space at
the end of the acquisition of each slice (i.e., reverse centric phase
encoding) is demonstrated to be a simple and robust method for
significantly reducing the relative saturation in all imaging planes.
View ordering and resolution dependence were confirmed in multiplanar
abdominal images. The added importance of reducing the artifact in
accelerated acquisition techniques (e.g., parallel imaging) is
particularly notable in multiplanar balanced steady state free
precession images in the brain. Magn Reson Med 63:1415-1421, 2010. c
Wiley-Liss, Inc.

#e1#

#s1#
#abstract1477.txt#


With the number of receivers available on clinical MRI systems now
ranging from 8 to 32 channels, data compression methods are being
explored to lessen the demands on the computer for data handling and
processing. Although software-based methods of compression after
reception lessen computational requirements, a hardware-based method
before the receiver also reduces the number of receive channels
required. An eight-channel Eigencoil array is constructed by placing a
hardware radiofrequency signal combiner inline after preamplification,
before the receiver system. The Eigencoil array produces
signal-to-noise ratio (

#e1#

#s1#
#abstract1478.txt#


Parallel and perpendicular diffusion properties of water in the rat
spinal cord were investigated 3 and 30 days after dorsal root axotomy,
a specific insult resulting in early axonal degeneration followed by
later myelin damage in the dorsal column white matter. Results from /i
q/-space analysis (i.e., the diffusion probability density function)
obtained with strong diffusion weighting were compared to conventional
anisotropy and diffusivity measurements at low b-values, as well as to
histology for axon and myelin damage. /i q/-Space contrasts included
the height (return to zero displacement probability), full width at
half maximum, root mean square displacement, and kurtosis excess of the
probability density function, which quantifies the deviation from
gaussian diffusion. Following axotomy, a significant increase in
perpendicular diffusion (with decreased kurtosis excess) and decrease
in parallel diffusion (with increased kurtosis excess) were found in
lesions relative to uninjured white matter. Notably, a significant
change in abnormal parallel diffusion was detected from 3 to 30 days
with full width at half maximum, but not with conventional diffusivity.
Also, directional full width at half maximum and root mean square
displacement measurements exhibited different sensitivities to white
matter damage. When compared to histology, the increase in
perpendicular diffusion was not specific to demyelination, whereas
combined reduced parallel diffusion and increased perpendicular
diffusion was associated with axon damage. Magn Reson Med 63:1323-1335,
2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract1479.txt#


Parallel radio frequency transmission has recently been explored as a
means of tailoring the spatial response of MR excitation. In
particular, parallel transmission is increasingly used to accelerate
radio frequency pulses that rely on time-varying gradient fields to
achieve selectivity in multiple dimensions. The design of the
underlying multiple-channel radio frequency waveforms is mostly based
on regularized least-squares optimization in close analogy with image
reconstruction in parallel imaging. However, this analogy has important
limitations. Unlike image reconstruction, the design of radio frequency
waveforms is subject to multiple strict constraints, which arise from
technical power limits, as well as safety limits on local and global
energy deposition in vivo. To optimize excitation profiles under such
strict constraints, it is proposed to depart from the regularization
strategy and rely on semidefinite programming instead. To render this
approach fast, it is performed in a reduced search space, which is
obtained by initial Lanczos iteration. The proposed algorithm is
demonstrated to enable efficient pulse optimization within exactly the
given constraints, including local specific absorption rate limits for
multiple compartments. It is also shown that the proposed approach
readily accommodates advanced forward models of the excitation process,
including the effects of local off-resonance and transverse relaxation.
Magn Reson Med 63:1280-1291, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract1482.txt#


The following topics are dealt with: parallel processing; hybrid
parallel algorithm; multicore cluster nodes; multiprocessor computers;
service-oriented grid computing; object-oriented OpenMP programming and
cell processor cluster.

#e1#

#s1#
#abstract1483.txt#


We report efficient implementation techniques for FFT-based dense
multivariate polynomial arithmetic over finite fields, targeting
multi-cores. We have extended a preliminary study dedicated to
polynomial multiplication and obtained a complete set of efficient
parallel routines in Cilk++ for polynomial arithmetic such as normal
form computation. Since bivariate multiplication applied to balanced
data is a good kernel for these routines, we provide an in-depth study
on the performance and the cut-off criteria of our different
implementations for this operation. We also show that, not only
optimized parallel multiplication can improve the performance of
higher-level algorithms such as normal form computation but also this
composition is necessary for parallel normal form computation to reach
peak performance on a variety of problems that we have tested.

#e1#

#s1#
#abstract1484.txt#


We present implementations of two data-mining algorithms on a CELL
processor, and on a low-cost CBEA (CELL Broadband Engine Architecture)
cluster using multiple PlayStation3 consoles. Typical batch-processing
environments are often unsuitable for interactive data-mining processes
that require repeated adjustments to parameters, preprocessing steps,
and data, while contemporary desktops do not offer sufficient resources
for the large datasets available today. Our implementations for the k
Nearest Neighbour algorithm and the Decision Tree scale linearly with
the number of samples in the training data and the number of
processors, and demonstrate runtimes of under a minute for up to 500
000 samples.

#e1#

#s1#
#abstract1485.txt#


Simulations of the mechanics of the left ventricle of the heart with
fluid-structure interaction benefit greatly from the parallel
processing power of a high performance computing cluster, such as
HPCVL. The objective of this paper is to describe the computational
requirements for our simulations. Results of parallelization studies
show that, as expected, increasing the number of threads per job
reduces the total wall clock time for the simulations. Further, the
speed-up factor increases with increasing problem size. Comparative
simulations with different computational meshes and time steps show
that our numerical solutions are nearly independent of the mesh density
in the solid wall (myocardium) and the time step duration. The results
of these tests allow our simulations to continue with the confidence
that we are optimizing our computational resources while minimizing
errors due to choices in spatial or temporal resolution.

#e1#

#s1#
#abstract1486.txt#


Modern computers operate at enormous speeds-capable of executing in
excess of 10/sup 13/ instructions per second-but their sequential
approach to processing, by which logical operations are performed one
after another, has remained unchanged since the 1950s. In contrast,
although individual neurons of the human brain fire at around just
10/sup 3/ times per second, the simultaneous collective action of
millions of neurons enables them to complete certain tasks more
efficiently than even the fastest supercomputer. Here we demonstrate an
assembly of molecular switches that simultaneously interact to perform
a variety of computational tasks including conventional digital logic,
calculating Voronoi diagrams, and simulating natural phenomena such as
heat diffusion and cancer growth. As well as representing a conceptual
shift from serial-processing with static architectures, our parallel,
dynamically reconfigurable approach could provide a means to solve
otherwise intractable computational problems.

#e1#

#s1#
#abstract1490.txt#


Malware analysis involves processing large amounts of storage to look
for suspicious files. This is time consuming and requires a large
amount of processing power, often affecting other applications running
on a personal computer. By using hardware included in most personal
computers, a performance increase can be seen. The processing of files
can be done on a graphics processing unit (GPU), contained in common
video cards. A GPU is perfect for this because of its strong
similarities to a central processing unit (CPU). Since the GPU has
multiple processing units (32-128) it has an advantage over a CPU (1-8
processing units). This allows a single data stream to be processed
using different metrics in parallel, while consuming minimal clock
cycles from the CPU. A GPU also has its own memory, separated from the
CPU, but also has the ability to share part of the CPU's memory. By
using the GPU on a personal computer, other applications and the file
processing for malware do not experience performance decreases by
fighting over CPU usage. Using a GPU for file processing can also be
applied to network intrusion detection systems (NIDS) and firewall
based applications in addition to anti-malware applications. This
research in progress investigates the use of GPUs for
monitoring/detecting malicious activity. Specifically, we are
processing files for malicious activity using different file sizes as
well as malicious and non-malicious files which allows the performance
and feasibility of using GPUs to be determined. We anticipate faster
anti-malware products, faster NIDS response times, faster firewall
applications, and a decrease in the time required to analyze files to
generate signatures and understand exactly what the file is doing.
UT INSPEC:11296895

#e1#

#s1#
#abstract1495.txt#


Cellular automata models have historically been a major approach to
studying the information-processing properties of self-replication.
Here we explore the feasibility of adopting genetic programming so
that, when it is given a fairly arbitrary initial cellular automata
configuration, it will automatically generate a set of rules that make
the given configuration replicate. We found that this approach works
surprisingly effectively for structures as large as 50 components or
more. The replication mechanisms discovered by genetic programming work
quite differently than those of many past manually designed
replicators: There is no identifiable instruction sequence or
construction arm, the replicating structures generally translate and
rotate as they reproduce, and they divide via a fissionlike process
that involves highly parallel operations. This makes replication very
fast, and one cannot identify which descendant is the parent and which
is the child. The ability to automatically generate self-replicating
structures in this fashion allowed us to examine the resulting
replicators as their properties were systematically varied. Further, it
proved possible to produce replicators that simultaneously deposited
secondary structures while replicating, as in some past manually
designed models. We conclude that genetic programming is a powerful
tool for studying self-replication that might also be profitably used
in contexts other than cellular spaces.

#e1#

#s1#
#abstract1496.txt#


The very long baseline interferometry (VLBI) technique currently
demands data storage resources of about 2 TBytes per day which must be
analyzed in correlation centers. The current process involves the
physical storing and shipping of magnetic disks to these centers, where
up to 5 days are needed for transportation. The need of a fast
turnaround opened a new research line where all the collected data of
observing radiotelescope stations is sent over the Internet. This
technique is called eVLBI. Ideally the station is part of a National
Research and Education Network (NREN) where multiple intercontinental
routes are available. Under this scenario a new protocol has been
developed which allows multiple parallel data flows with important
throughput improvements. The unique properties of VLBI data imply the
development of a custom load control based on user datagram protocol
(UDP). A description of the new protocol and performance comparisons of
the first demonstrations for eVLBI performed at the Transportable
Integrated Geodetic Observatory (TIGO) are included in the present
article. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1497.txt#


Drainage networks determination from digital elevation models (DEM) has
been a widely studied problem in the last three decades. During this
time, satellite technology has been improving and optimizing
digitalized images, and computers have been increasing their
capabilities to manage such a huge quantity of information. The rapid
growth of CPU power and memory size has concentrated the discussion of
DEM algorithms on the accuracy of their results more than their running
times. However, obtaining improved running times remains crucial when
DEM dimensions and their resolutions increase. Parallel computation
provides an opportunity to reduce run times. Recently developed
graphics processing units (GPUs) are computationally fast not only in
Computer Graphics but in General Purpose Computation, the so-called
GPGPU. In this paper we explore the parallel characteristics of these
GPUs for drainage network determination, using the C-oriented language
of CUDA developed by NVIDIA. The results are simple algorithms that run
on low-cost technology with a high performance response, obtaining CPU
improvements of up to 8*. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1499.txt#


This paper presents an ontology-based approach for the design of a
collaborative business process model (CBP). This CBP is considered as a
specification of needs in order to build a collaboration information
system (CIS) for a network of organizations. The study is a part of a
model-driven engineering approach of the CIS in a specific enterprise
interoperability framework that will be summarised. An adaptation of
the Business Process Modelling Notation (BPMN) is used to represent the
CBP model. We develop a knowledge-based system (KbS) which is composed
of three main parts: knowledge gathering, knowledge representation and
reasoning, and collaborative business process modelling. The first part
starts from a high abstraction level where knowledge from business
partners is captured. A collaboration ontology is defined in order to
provide a structure to store and use the knowledge captured. In
parallel, we try to reuse generic existing knowledge about business
processes from the MIT Process Handbook repository. This results in a
collaboration process ontology that is also described. A set of rules
is defined in order to extract knowledge about fragments of the CBP
model from the two previous ontologies. These fragments are finally
assembled in the third part of the KbS. A prototype of the KbS has been
developed in order to implement and support this approach. The
prototype is a computer-aided design tool of the CBP. In this paper, we
will present the theoretical aspects of each part of this KbS as well
as the tools that we developed and used in order to support its
functionalities. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1514.txt#


The scheduling and mapping of the precedence-constrained task graph to
processors is considered to be the most crucial NP-complete problem in
parallel and distributed computing systems. Several genetic algorithms
have been developed to solve this problem. A common feature in most of
them has been the use of chromosomal representation for a schedule.
However, these algorithms are monolithic, as they attempt to scan the
entire solution space without considering how to reduce the complexity
of the optimization process. In this paper, two genetic algorithms have
been developed and implemented. Our developed algorithms are genetic
algorithms with some heuristic principles that have been added to
improve the performance. According to the first developed genetic
algorithm, two fitness functions have been applied one after the other.
The first fitness function is concerned with minimizing the total
execution time (schedule length), and the second one is concerned with
the load balance satisfaction. The second developed genetic algorithm
is based on a task duplication technique to overcome the communication
overhead. Our proposed algorithms have been implemented and evaluated
using benchmarks. According to the evolved results, it has been found
that our algorithms always outperform the traditional algorithms. [All
rights reserved Elsevier].

#e1#

#s1#
#abstract1515.txt#


The paper presents algorithmic solutions dedicated to computer
navigation system which is to assist bronchoscope positioning during
transbronchial needle-aspiration biopsy. The navigation exploits the
principle of on-line registration of real images from endoscope camera
and the virtual ones generated on the base of computed tomography (CT)
data of a patient. When these images are similar an assumption is made,
that the bronchoscope and virtual camera have approximately the same
position and view direction. In this paper the following computational
aspects are described: correction of camera lens distortion, fast
approximate estimation of endoscope ego-motion, reconstruction of
bronchial tree from the CT data by means of their segmentation and its
centerline calculation, virtual views generation, registration of real
and virtual images via maximization of their mutual information and,
finally, efficient parallel and network implementation of the
navigation system which is under development at present. [All rights
reserved Elsevier].

#e1#

#s1#
#abstract1523.txt#


The steam reforming of methane in a parallel plate microreactor,
consisting of alternating channels carrying out catalytic combustion
and reforming on opposite sides of a wall, is modeled with fundamental
kinetics and a pseudo-2D reactor model. It is shown that at high fuel
conversions, the choice of hydrocarbon combustible fuel is immaterial
when suitable compositions are used so that the energy input is kept
the same. On the other hand, direct comparison of Rh and Ni indicates
that the choice of reforming catalyst is critical. Speed up of heat
transfer via miniaturization is insufficient for process
intensification; catalyst-intensification is also needed to avoid hot
spots and enable compact devices for portable and distributed power
generation. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1525.txt#


A block-oriented nonlinear system is composed of a concatenation of
blocks representing either memoryless nonlinearities or linear dynamic
subsystems. Wiener, Hammerstein, and Wiener-Hammerstein (WH) models are
the most commonly used ones for modeling such block-oriented nonlinear
systems. In this paper, we develop a tensor analysis-based approach for
determining both the structure and the parameters of the most
appropriate model among the three above listed models. The structure is
deduced from the rank of a tensor obtained by convolving a random
finite impulse response (FIR) linear filter with a /i p/ th-order
Volterra kernel (/i p/ # 2), associated with the block-oriented
nonlinear system to be identified. The parameters of the linear
subsystems are obtained from the PARAllel FACtor analysis (PARAFAC)
decomposition of the /i p/th-order Volterra kernel associated with the
original nonlinear system and/or an extended WH system, whereas those
of the nonlinear subsystem are estimated using the least squares
method. The performance of the proposed identification scheme is
illustrated by means of some simulation results.

#e1#

#s1#
#abstract1526.txt#


In this paper, we study the problem of joint model selection and
parameter estimation under the Bayesian framework. We propose to use
the Population Monte Carlo (PMC) methodology in carrying out Bayesian
computations. The PMC methodology has recently been proposed as an
efficient sampling technique and an alternative to Markov Chain Monte
Carlo (MCMC) sampling. Its flexibility in constructing transition
kernels allows for joint sampling of parameter spaces that belong to
different models. The proposed method is able to estimate the desired
/i a posteriori/ distributions accurately. In comparison to the
Reversible Jump MCMC (RJMCMC) algorithm, which is popular in solving
the same problem, the PMC algorithm does not require burn-in period, it
produces approximately uncorrelated samples, and it can be implemented
in a parallel fashion. We demonstrate our approach on two examples:
sinusoids in white Gaussian noise and direction of arrival (DOA)
estimation in colored Gaussian noise, where in both cases the number of
signals in the data is /i a priori/ unknown. Both simulations show the
effectiveness of our proposed algorithm.

#e1#

#s1#
#abstract1527.txt#


The convergence of integrated electronic devices with nanotechnology
structures on heterogeneous systems presents promising opportunities
for the development of new classes of rapid, sensitive, and reliable
sensors. The main advantage of embedding microelectronic readout
structures with sensing elements is twofold. On the one hand, the

#e1#

#s1#
#abstract1529.txt#


This paper deals with solving large instances of the Linear Sum
Assignment Problems (LSAPs) under realtime constraints, using Graphical
Processing Units (GPUs). The motivating scenario is an industrial
application for P2P live streaming that is moderated by a central
tracker that is periodically solving LSAP instances to optimize the
connectivity of thousands of peers. However, our findings are generic
enough to be applied in other contexts. Our main contribution is a
parallel version of a heuristic algorithm called Deep Greedy Switching
(DGS) on GPUs using the CUDA programming language. DGS sacrifices
absolute optimality in favor of a substantial speedup in comparison to
classical LSAP solvers like the Hungarian and auctioning methods. We
show the modifications needed to parallelize the DGS algorithm and the
performance gains of our approach compared to a sequential CPU-based
implementation of DGS and a mixed CPU/GPU-based implementation of it.

#e1#

#s1#
#abstract1530.txt#


The multi-objective integer programming problems in large scale are
considered time consuming. In the past, mathematical structures were
used that can get benefits of high processing powers and parallel
processing. The Branch and Bound (B&B) algorithm is one of the most
used methods to solve combinatorial optimization problems. A general
approach to generate all non-dominated solutions of the multi-objective
integer programming (MOIP) Problem is developed. In this paper, a
supervisor-master-sub-master-worker algorithm to solve large scale
integer multi-objective problems that can get the benefits of
mathematical structures and high processing powers has been proposed.
This approach addresses several issues related to the characteristics
of the algorithm itself and the properties of parallel computing
systems. From the solved benchmark example this algorithm proved to
provide a considerable high performance. Results show that a
consistently better efficiency can be achieved in solving integer
equations, providing reduction of time.

#e1#

#s1#
#abstract1531.txt#


NLP search takes a long amount of time due to large size of corpus,
besides there are too many hits at the server. At present, the strategy
to deal with search engines is to have many thousands of servers in
order to provide real-time searches. Fast alternatives are therefore
sought. In this paper, we present a pioneering work in this direction
by taking word stemming, a crucial aspect of search and indexing
algorithms and showing how significant performance gain can be
accomplished by employing multi-core architectures, which will serve
the purpose of home computers in near future. We present our analysis
of Porter's stemming algorithm on Cell Broadband Engine and describe
the manner in which SIMD operations can be utilized to maximize
performance. Our results show that cell processors provide performance
gains of over 50 times over popular Intel processors and hence possess
tremendous potential for NLP-IR applications.

#e1#

#s1#
#abstract1532.txt#


The scheduling and mapping of the precedence-constrained task graphs of
parallel programs to processors is considered one of the most crucial
NP-complete problems in parallel and distributed computing systems. In
this paper, a dynamic task scheduling model based on fuzzy logic is
proposed. The main objective of this technique is to improve the fuzzy
decision which is used in task scheduling on a network of processing
elements by introducing new input parameters to an existing fuzzy model
and, in the same time, improving the load balance on the network in a
dynamic environment. The proposed Fuzzy Model is capable of processing
inputs from on the fly data that arises from the current state of the
processors. According to the proposed model, tasks are generated
randomly and are served based on the First-Come-First-Serve rule. When
the task is ready to be assigned, its information is passed to the
processors for bidding. Each processor has a local scheduler for
managing its own activities, which supplies information on its current
state and follows whatever decision is given where the fuzzy logic
mechanism is used in making decision on the task assignment. A
comparative study between the existed fuzzy model and our modified
fuzzy model has been done. The comparative results show that our
modified fuzzy model outperforms the existed one.

#e1#

#s1#
#abstract1533.txt#


Many computing-intensive applications in different domains are
benefiting from new High Performance Computing (HPC) architectures such
as clusters and computational grids. This experimental study shows the
performance results of parallelizing a computing-intensive risk
management financial simulator application using the Message Passing
Interface (MPI) and running it on two configurations of a dedicated
grid consisting of workstations on three geographically distant
locations. To benchmark the grid configurations performance, the
application was also run on a local cluster. Results showed that gained
speed ups are scalable and those of the two grid configurations are
close to the local cluster suggesting that more speed up gains can be
achieved with grid expansion.

#e1#

#s1#
#abstract1534.txt#


Directly from the surface coil images, a new adaptive smoothing method
for estimating the surface coil sensitivity is proposed in this paper.
When the coil sensitivity maps estimated by this method are applied to
reconstruct the full Field-Of-View image from the under-sampled images,
the quality of the reconstructed image in parallel magnetic resonance
imaging is improved. As a result, without using the body coil for
additional reference scans, the proposed method is valuable for
determining the coil sensitivity maps alone from the receive coil
surface images. Furthermore, the study innovatively poses that the
relatively distinct coil sensitivity maps could be utilized to
reconstruct full Field-Of-View image in aim of speed-up magnetic
resonance imaging.

#e1#

#s1#
#abstract1535.txt#


In recent years, hardware based packet classification has became an
essential component in many networking devices. Ternary
Content-Addressable Memories (TCAMs) are one of the most popular
solutions in this domain, allowing to compare in parallel the packet
header against a large set of rules, and to retrieve the first match.
However, using TCAM to match a range of values is much more problematic
and dramatically reduces the cost effectiveness of the solution. In
this paper we study ways to use simple built-in TCAM mechanisms in
order to increase the efficiency of range coverage. While current
techniques have a worst expansion ratio of 2W-4, we present an
efficient algorithm enabling to encode any range with at most W TCAM
entries (where W in the number of bits), without using additional
processing, extra bits, and without any external encoding. The same
paradigm can be applied to multiple raging rules as well, resulting in
significant improvement over current known techniques. Moreover, our
simulation results indicate that these techniques can be used to reduce
the actual TCAM size of hardware networking devices under realistic
scenarios.

#e1#

#s1#
#abstract1536.txt#


In this paper, we identify the unique challenges in deploying
parallelism on TCAM-based pattern matching for Network Intrusion
Detection Systems (NIDSes). We resolve two critical issues when
designing scalable parallelism specifically for pattern matching
modules: 1) how to enable fine-grained parallelism in pursuit of
effective load balancing and desirable speedup simultaneously; and 2)
how to reconcile the tension between parallel processing speedup and
prohibitive TCAM power consumption. To this end, we first propose the
novel concept of Negative Pattern Matching to partition flows, by which
the number of TCAM lookups can be significantly reduced, and the
resulting (fine-grained) flow segments can be inspected in parallel
without incurring false negatives. Then we propose the notion of
Exclusive Pattern Matching to divide the entire pattern set into
multiple subsets which can later be matched against selectively and
independently without affecting the correctness. We show that Exclusive
Pattern Matching enables the adoption of smaller and faster TCAM blocks
and improves both the pattern matching speed and scalability. Finally,
our theoretical and experimental results validate that the above two
concepts are inherently complementary, enabling our integrated scheme
to provide performance gain in any scenario (with either clean or dirty
traffic).

#e1#

#s1#
#abstract1537.txt#


From the constitution of computation, the main subjects of computer
science and technology have been summarized. Meanwhile, the reasons why
the developments of software technologies have pushed the computational
science be widely used were analyzed. Multiple computing models were
discussed base on different computer architectures. Finally,
computation as a methodology, two important finance problems, effective
financial supervision and high-speed financial product pricing, were
solved by distributed and parallel computing.

#e1#

#s1#
#abstract1538.txt#


Forward-projection is an important process of computed tomography
reconstruction. An accurate forward-projection has a significant impact
on increasing the reconstructed quality of iterative algorithms.
However its computations are very costly. In order to improve
performance of forward-projection, we proposed a GPU-accelerated scheme
for it with a new sampling method. There are two novel techniques in
our accelerating scheme. One is parallel strategy based on view angles
to project volume at different view angles simultaneously, the other is
distance-based sampling manner in volume to make full use of texture
cache. A number of experiments were performed to verify the proposed
algorithm and the results show that this method is accurate and
efficient.

#e1#

#s1#
#abstract1539.txt#


This paper proposes that multifractal analysis method be adopted for
the analysis of edge images detected with classical edge detectors. We
analyze the direct multifractal analysis method based on box-counting
method with experimental images, and determine the value of f(q) and
a(q). By using this method, we find that the edge images detected by
difference edge detectors are similar, but, the multifractal
micro-structures of them are different, and, that of edge images
detected by Prewitt and Roberts detector are parallel, and so do that
by Log and Zerocross detector.

#e1#

#s1#
#abstract1540.txt#


All-printed electronics as a means of achieving ultra-low-cost
electronic circuits has attracted great interest in recent years.
Inkjet printing is one of the most promising techniques by which the
circuit components can be ultimately drawn (i.e. printed) onto the
substrate in one step. Here, the inkjet printing technique was used to
chemically deposit silver nanoparticles (10-200 nm) simply by ejection
of silver nitrate and reducing solutions onto different substrates such
as paper, PET plastic film and textile fabrics. The silver patterns
were tested for their functionality to work as circuit components like
conductor, resistor, capacitor and inductor. Different levels of
conductivity were achieved simply by changing the printing sequence,
inks ratio and concentration. The highest level of conductivity
achieved by an office thermal inkjet printer (300 dpi) was 5.54 *
10/sup 5/ Sm/sup -1/ on paper. Inkjet deposited capacitors could
exhibit a capacitance of more than 1.5 nF (parallel plate 45 * 45
mm/sup 2/) and induction coils displayed an inductance of around 400 mu
H (planar coil 10 cm in diameter). Comparison of electronic performance
of inkjet deposited components to the performance of conventionally
etched items makes the technique highly promising for fabricating
different printed electronic devices.

#e1#

#s1#
#abstract1541.txt#


A hybrid parallel and out-of-core algorithm pads blocks from a
structured grid with layers of ghost data from adjacent blocks. This
enables end-to-end streaming computations on very large data sets that
gracefully adapt to available computing resources, from a
single-processor machine to parallel visualization clusters.

#e1#

#s1#
#abstract1542.txt#


Supporting Distributed Shared Memory (DSM) is essential for multi-core
Network-on-Chips for the sake of reusing huge amount of legacy code and
easy programmability. We propose a microcoded controller as a hardware
module in each node to connect the core, the local memory and the
network. The controller is programmable where the DSM functions such as
virtual-to-physical address translation, memory access and
synchronization etc. are realized using microcode. To enable concurrent
processing of memory requests from the local and remote cores, our
controller features two mini-processors, one dealing with requests from
the local core and the other from remote cores. Synthesis results
suggest that the controller consumes 51k gates for the logic and can
run up to 455 MHz in 130 nm technology. To evaluate its performance, we
use synthetic and application workloads. Results show that, when the
system size is scaled up, the delay overhead incurred by the controller
may become less significant when compared with the network delay. In
this way, the delay efficiency of our DSM solution is close to hardware
solutions on average but still have all the flexibility of software
solutions.

#e1#

#s1#
#abstract1543.txt#


Wearable, mobile computing platforms are envisioned to be used in
out-patient monitoring and care. These systems continuously perform
signal filtering, transformations, and classification, which are quite
compute intensive, and quickly drain the system energy. The design
space of these human activity sensors is large and choice of sampling
frequency, feature detection algorithm, length of the window of
transition detection etc., and all these choices fundamentally
trade-off power/performance for accuracy of detection. In this work, we
explore this design space, and make several interesting conclusions
that can be used as rules of thumb for quick, yet power-efficient
designs of such systems. For instance, we find that the x-axis of our
signal, which was oriented to be parallel to the forearm, is the most
important signal to be monitored, for our set of hand activities. Our
experimental results show that by carefully choosing system design
parameters, there is considerable (5X) scope of improving the
performance/power of the system, for minimal (5%) loss in accuracy.

#e1#

#s1#
#abstract1544.txt#


This paper presents a novel hardware architecture for genetic vector
quantizer (VQ) design. It is based on steady-state genetic algorithm
(GA) and adopts shift registers for accelerating mutation and crossover
operations while reducing area cost. It also uses a pipeline for
fitness evaluation. The proposed architecture has been embedded in a
softcore CPU for physical performance measurement. Experimental results
show that it is an effective alternative for VQ optimization attaining
both high performance and low computational time. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract1545.txt#


We investigated hydrogenated aluminum oxide (a-Al/sub 1-x/O/sub x/:H)
as a high quality rear surface passivation layer of crystalline silicon
solar cells. The a-Al/sub 1-x/O/sub x/:H films were deposited by
plasma-enhanced chemical vapor deposition (PECVD) using a mixture of
trimethylaluminum (TMA), carbon dioxide (

#e1#

#s1#
#abstract1546.txt#


This paper discusses whether optical technology could offer a solution
to the heat generation and bandwidth limitations that the computing
industry is starting to face? The benefits of energy-efficient passive
components, low crosstalk and parallel processing suggest that the
answer may be yes.

#e1#

#s1#
#abstract1558.txt#


Current state-of-the-art task scheduling algorithms for network packet
processing schedule the program into a parallel-pipeline topology on
network processors to maximize the throughput. However, there has been
no existing work targeting power budget for packet processing on
off-the-shelf multicore architectures. As energy consumption,
reliability and cooling cost for packet processing systems become
increasingly important, it is necessary to integrate power-awareness
into a scheduler to meet the power budget. In this paper, we propose a
novel scheduling algorithm to optimize both throughput and latency
given a power budget for network packet processing on multicore
architectures. This algorithm addresses power-aware parallel-pipeline
scheduling problem by applying per-core DVFS to optimally adjust
frequency on each core. We implement our algorithm on an AMD machine
with two Quad-Core Opteron 2350 processors and compare the results with
existing algorithms given the same power budget. For six real packet
processing applications, our algorithm improves throughput and reduces
latency by an average of 64.6% and 25.2%, respectively.

#e1#

#s1#
#abstract1559.txt#


The data exchange capacity among each processing units was a key factor
in performance of multiprocessor parallel system. Parallel
interconnection based on master-slave architecture constrained by
several factors could not be further increased bandwidth. However,
high-speed serial interconnect based on point-to-point architecture
could achieve greater bandwidth. This paper presented a new
multi-processor parallel system based on high-speed serial
interconnects. There were four processing channels in this system. Each
channel consisted of a piece of FPGA and a piece of DSP. All of these
processing units were connected by serial RapidIO bus. Several data
exchange methods were designed in this project to solve the bottleneck
of data exchange in the multi-processor parallel system. Therefore, the
system parallelism and performance was greatly improved. The system has
some characteristic: versatility, high parallelism, structure flexible,
modularization.

#e1#

#s1#
#abstract1560.txt#


A distributed monitoring and diagnosis system to complex
electromechanical equipments is researched and outlined. It includes
three layers: monitoring units, bus communication and information
processing and application system and it is constructed via BITBUS
field bus and distributed network. Its software modules, including
communication, monitoring, signal processing and fault diagnosis, have
been designed.

#e1#

#s1#
#abstract1561.txt#


Rectification of stereo vision is also call rectification of Epipolar
Lines. For searching the corresponding points rapidly and accurately,
the rectification is used to make the epipolar lines of stereo images
be parallel to the horizontal direction and remove the parallax in
vertical direction. In the paper, a robust algorithm is presented to
rectify the stereo images based on the traditional projective
rectification algorithm. In this method, the rectification is composed
of several steps, and the transformation matrix is calculated by the
corresponding points and then is optimized by Levenberg-Marquardt
method and Evolutionary Programming (EP) algorithm. And meanwhile the
RANSAC algorithm is used, which avoid the severe error caused by the
noise and wrong correspondence. Furthermore, a robust optimization
function is proposed for the situation of slight movement of two
cameras. Experiments with real images show that the method is an
effective rectification method.

#e1#

#s1#
#abstract1562.txt#


Deep Brain Stimulation (DBS) has been successfully used throughout the
world for the treatment of Parkinson's disease symptoms. To control
abnormal spontaneous electrical activity in target brain areas DBS
utilizes a continuous stimulation signal. This continuous power draw
means that its implanted battery power source needs to be replaced
every 18-24 months. To prolong the life span of the battery, a
technique to accurately recognize and predict the onset of the
Parkinson's disease tremors in human subjects and thus implement an
on-demand stimulator is discussed here. The approach is to use a radial
basis function neural network (RBFNN) based on particle swarm
optimization (PSO) and principal component analysis (PCA) with Local
Field Potential (LFP) data recorded via the stimulation electrodes to
predict activity related to tremor onset. To test this approach, LFPs
from the subthalamic nucleus (STN) obtained through deep brain
electrodes implanted in a Parkinson patient are used to train the
network. To validate the network's performance, electromyographic (EMG)
signals from the patient's forearm are recorded in parallel with the
LFPs to accurately determine occurrences of tremor, and these are
compared to the performance of the network. It has been found that
detection accuracies of up to 89% are possible. Performance comparisons
have also been made between a conventional RBFNN and an RBFNN based on
PSO which show a marginal decrease in performance but with notable
reduction in computational overhead.

#e1#

#s1#
#abstract1563.txt#


The noise power spectrum (NPS) is a useful metric for understanding the
noise content in images. To examine some unique properties of the NPS
of fan beam CT, the authors derived an analytical expression for the
NPS of fan beam CT and validated it with computer simulations. The
nonstationary noise behavior of fan beam CT was examined by analyzing
local regions and the entire field-of-view (FOV). This was performed
for cases with uniform as well as nonuniform noise across the detector
cells and across views. The simulated NPS from the entire FOV and local
regions showed good agreement with the analytically derived NPS. The
analysis shows that whereas the NPS of a large FOV in parallel beam CT
(using a ramp filter) is proportional to frequency, the NPS with direct
fan beam FBP reconstruction shows a high frequency roll off. Even in
small regions, the fan beam NPS can show a sharp transition
(discontinuity) at high frequencies. These effects are due to the
variable magnification and therefore are more pronounced as the fan
angle increases. For cases with nonuniform noise, the NPS can show the
directional dependence and additional effects.

#e1#

#s1#
#abstract1564.txt#


This paper presents a new algorithm for smoothing 3D binary images in a
topology preserving way. Our algorithm is a reduction operator: some
border points that are considered as extremities are removed. The
proposed method is composed of two parallel reduction operators. We are
to apply our smoothing algorithm as an iteration-by-iteration pruning
for reducing the noise sensitivity of 3D parallel surface-thinning
algorithms. An efficient implementation of our algorithm is sketched
and its topological correctness for (26,6) pictures is proved.

#e1#

#s1#
#abstract1565.txt#


We observe that two-dimensional array languages are useful in Image
Processing and Analysis. We represent two-dimensional iso-picture
languages in terms of various classes of polyoisominoes. Polyoisominoes
are related to tiling. The art of tiling has played an important role
in the field of architecture since early civilization. A parallel
generating model called pasting system was proposed in order to form
coloured patterns over square grid (1999). This model allows two cells
(tiles) to get glued on their sides depending upon some rules. Two
tiles are called adjacent if they have an edge in common. We prove that
various classes of polyoisominoes denned are tiling recognizable
languages and local language. Local languages are introduced by
Giammarresi and Restivo (1997).

#e1#

#s1#
#abstract1566.txt#


This paper presents a new type of inductive power transfer (IPT) pickup
that directly regulates the power in ac form, hence producing a
controllable high-frequency ac source suitable for lighting
applications. The pickup has significant advantages in terms of
increasing system efficiency, reducing pickup size, and lowering
production cost compared to traditional pickups that also produce a
controlled ac output using complex ac-dc-ac conversion circuits. The
new ac processing pickup employs switches operating under
zero-voltage-switching conditions to clamp parts of the resonant
voltage across a parallel tuned /i LC/ resonant tank to achieve power
regulation over a wide load range. The operation of the pickup is
analyzed and the circuit waveforms have been verified by experimental
results. A complete IPT system using the ac processing pickup was
tested on a 500-W lighting system and an efficiency of 96% was obtained
when delivering 500 W to multiple resistive light bulbs.

#e1#

#s1#
#abstract1580.txt#


Non-chemically treated single-walled carbon nanotubes (SWCNTs) were
assembled at a liquid-liquid interface in the shape of an ultrathin
film. SWCNTs were dispersed into water phase by aid of sodium dodecyl
sulfate (SDS) and were assembled at a hexane-water interface with
addition of ethanol. The assembled film was transferred onto a solid
substrate at various dipping speeds. Polarized absorption spectra of
the transferred film indicated that the SWCNTs aligned with their long
axis parallel to the dipping direction when the assembly was
transferred at the dipping speed of 50 mm/min.

#e1#

#s1#
#abstract1581.txt#


The shared-memory model has been adopted, both for data exchange as
well as synchronization using semaphores in almost every on-chip
multiprocessor implementation, ranging from general purpose chip
multiprocessors (CMPs) to domain specific multi-core graphics
processing units (GPUs). Low-latency synchronization is desirable but
is hard to achieve in practice due to the memory hierarchy. On the
contrary, an explicit exchange of synchronization tokens among the
processing elements through dedicated on-chip links would be beneficial
for the overall system performance. In this paper we propose the Medea
NoC-based framework, a hybrid shared-memory/message-passing approach.
Medea has been modeled with a fast, cycle-accurate SystemC
implementation enabling a fast system exploration varying several
parameters like number and types of cores, cache size and policy and
NoC features. In addition, every SystemC block has its RTL counterpart
for physical implementation on FPGAs and ASICs. A parallel version of
the Jacobi algorithm has been used as a test application to validate
the methodology. Results confirm expectations about performance and
effectiveness of system exploration and design.

#e1#

#s1#
#abstract1582.txt#


Throughput and programmability have always been the central, but
generally conflicting concerns for modern IP router designs. Current
high performance routers depend on proprietary hardware solutions,
which make it difficult to adapt to ever-changing network protocols. On
the other hand, software routers offer the best flexibility and
programmability, but could only achieve a throughput one order of
magnitude lower. Modern GPUs are offering significant computing power,
and its data-parallel computing model well matches the typical patterns
of packet processing on routers. Accordingly, in this research we
investigate the potential of CUDA-enabled GPUs for IP routing
applications. As a first step toward exploring the architecture of a
GPU based software router, we developed GPU solutions for a series of
core IP routing applications such as IP routing table lookup and
pattern match. For the deep packet inspection application, we
implemented both a Bloom-filter based string matching algorithm and a
finite automata based regular expression matching algorithm. A GPU
based routing table lookup solution is also proposed in this work.
Experimental results proved that GPU could accelerate the routing
processing by one order of magnitude. Our work suggests that, with
proper architectural modifications, GPU based software routers could
deliver significant higher throughput than previous CPU based solutions.

#e1#

#s1#
#abstract1583.txt#


Emerging TSV-based 3D integration technologies have shown great promise
to overcome scalability limitations in 2D designs by stacking multiple
memory dies on top of a many-core die. Application software developers
need programming models and tools to fully exploit the potential of
vertically stacked memory. In this work, we focus on efficient data
mapping for SPMD parallel applications on an explicitly managed
3D-stacked memory hierarchy, which requires placement of data across
multiple vertical memory stacks to be carefully optimized. We propose a
programming framework with compiler support that enables array
partitioning. Partitions are mapped to the 3D-stacked memory on top of
the processor that mostly accesses it to take advantage of the lower
latencies of vertical interconnect and for minimizing high-latency
traffic on the horizontal plane.

#e1#

#s1#
#abstract1584.txt#


Subdivision Surfaces provide a compact way to describe a smooth surface
using a mesh model. They are widely used in 3D animation and nearly all
modern modeling programs support them. In this work we describe a
complete parallel pipeline for real-time interactive editing,
processing and rendering of smooth surface primitives on the Cell BE.
Our approach makes it possible to edit and render these high-order
graphics primitives passing them directly to a parallel pipeline which
tessellates them just before rendering. We describe a combination of
algorithmic, architectural and back-end optimizations that enable us to
render smooth subdivision surfaces in real time and to dynamically
deform 3D models represented by subdivision surfaces.

#e1#

#s1#
#abstract1585.txt#


Face Recognition techniques are solutions used to quickly screen a huge
number of persons without being intrusive in open environments or to
substitute id cards in companies or research institutes. There are
several reasons that require to systems implementing these techniques
to be reliable. This paper presents the design of a reliable face
recognition system implemented on Field Programmable Gate Array (FPGA).
The proposed implementation uses the concepts of multiprocessor
architecture, parallel software and dynamic reconfiguration to satisfy
the requirement of a reliable system. The target multiprocessor
architecture is extended to support the dynamic reconfiguration of the
processing unit to provide reliability to processors fault. The
experimental results show that, due to the multiprocessor architecture,
the parallel face recognition algorithm can achieve a speed up of 63%
with respect to the sequential version. Results regarding the overhead
in maintaining a reliable architecture are also shown.

#e1#

#s1#
#abstract1586.txt#


Future embedded system products, e.g. smart hand-held mobile terminals,
will accommodate a large number of applications that will partly run
sequentially and independently, partly concurrently and interacting on
massively parallel computing platforms. Already for systems of moderate
complexity, the design space will be huge and its exploration requires
that the system architect is able to quickly evaluate the performances
of candidate architectures and application mappings. The mainstream
evaluation technique today is the system-level performance simulation
of the applications and platforms using abstracted workload and
processing capacity models, respectively. These virtual system models
allow fast simulation of large systems at an early phase of development
with reasonable modeling effort and time. The accuracy of the
performance results is dependent on how closely the models used reflect
the actual system. This paper presents a compiler based technique for
automatic generation of workload models for performance simulation,
while exploiting an overall approach and platform performance capacity
models developed previously. The resulting workload models are
experimented using x264 video and JPEG encoding application examples.

#e1#

#s1#
#abstract1587.txt#


Heterogeneous reconfigurable processing architectures are often limited
by the speed at which they can access data in external memory. Such
architectures are designed for flexibility to support a broad range of
target applications, including advanced algorithms with significant
processing and data requirements. Clearly, strong performance of
applications in this category is an extremely relevant metric for
demonstrating the full performance potential of heterogeneous computing
platforms. One such example, a film grain noise reduction application
for high-definition video, which is composed of multiple image
processing tasks, requires enormous data rates due to its large input
image size and real-time processing constraints. This application is
especially representative of highly parallel, heterogeneous,
data-intensive programs that can properly exploit the advantages
offered by computing platforms with multiple heterogeneous
reconfigurable processing elements. To accomplish this task and meet
the above requirements, a bandwidth-optimized external memory
controller has been designed for use with a heterogeneous
reconfigurable architecture and its NoC interconnect. With the help of
the application described above, this paper evaluates the proposed
architecture in two forms: (1) with a basic memory controller IP and
(2) with the advanced memory controller design. The results illustrate
the full potential of the computing platform as well as the power of
heterogeneous reconfigurable computing combined with high-speed access
to large external memories.

#e1#

#s1#
#abstract1588.txt#


Motion Estimation (ME) is the most computationally intensive part of
video compression and video enhancement systems. One bit transform
(1BT) based ME algorithms have low computational complexity. Therefore,
in this paper, we propose a high performance reconfigurable hardware
architecture of 1BT based multiple reference frame (MRF) ME. The
proposed ME hardware architecture performs full search ME for 4
Macroblocks and 4 reference frames in parallel. The proposed hardware
is faster than the 1BT based ME hardware reported in the literature
even though it is capable of searching in 4 reference frames. MRF ME
increases the ME performance at the expense of increased computational
complexity. The reconfigurability of the proposed ME hardware is used
to statically configure the number and selection of reference frames
based on the application requirements in order to trade-off ME
performance and computational complexity. The proposed hardware
architecture is implemented in Verilog HDL. The MRF ME hardware
consumes %65 of the slices in a Xilinx XC2VP30-7 FPGA. It can work at
191 MHz in the same FPGA and is capable of processing 83 1920 * 1080
full High Definition frames per second.

#e1#

#s1#
#abstract1589.txt#


Parallel file systems are very sensitive to adverse conditions, and the
lack of synergy between such file systems and some of the applications
running on them has a negative impact on the overall system
performance. Our observations indicate that the increased pressure on
metadata management is one of the relevant causes of performance drops.
This paper proposes a virtualization layer above the native file system
that, transparently to the user, reorganizes the underlying directory
tree, mitigating bottlenecks by taking advantage of the native file
system optimizations and limiting the effects of potentially harmful
application behavior. We developed

#e1#

#s1#
#abstract1590.txt#


Multi-Processor System-on-Chips (MPSoCs) exploit task-level parallelism
to achieve high computation throughput, but concurrent memory accesses
from multiple PEs may cause memory bottleneck. Therefore, to maximize
system performance, it is important to simultaneously consider the PE
and on-chip memory architecture design. However, in a traditional MPSoC
design flow, PE allocation and on-chip memory allocation are often
considered independently. To tackle this problem, we propose the first
PE and Memory Co-synthesis (PM-

#e1#

#s1#
#abstract1591.txt#


This paper summarizes a special session on multi-core/multi-processor
system-on-chip (MPSoC) programming challenges. Wireless multimedia
terminals are among the key drivers for MPSoC platform evolution.
Heterogeneous multi-processor architectures achieve high performance
and can lead to a significant reduction in energy consumption for this
class of applications. However, just designing energy efficient
hardware is not enough. Programming models and tools for efficient
MPSoC programming are equally important to ensure optimum platform
utilization. Unfortunately, this discipline is still in its infancy,
which endangers the return on investment for MPSoC architecture
designs. On one hand there is a need for maintaining and gradually
porting a large amount of legacy code to MPSoCs. On the other hand,
special C language extensions for parallel programming as well as
adapted process network programming models provide a great opportunity
to completely rethink the traditional sequential programming paradigm
for sake of higher efficiency and productivity. MPSoC programming is
more than just code parallelisation, though. Besides energy efficiency,
limited and specialized processing resources, and real-time constraints
also growing software complexity and mapping of simultaneous
applications need to be taken into account. We analyze the programming
methodology requirements for heterogeneous MPSoC platforms and outline
new approaches.

#e1#

#s1#
#abstract1592.txt#


In order to solve the challenges in processor design for the next
generation wireless communication systems, this paper first proposes a
system level design flow for communication domain specific processor,
and then proposes a novel processor architecture for the next
generation wireless communication named GAEA using this design flow.
GAEA is a shared memory multi-core SoC based on Software Controlled
Time Division Multiplexing Bus, with which programmers can easily
explore memory-level parallelism of applications by proper instructions
and scheduling algorithms. MPE, which is the kernel component of GAEA,
adopts hybrid parallel processing scheme to explore instruction-level
and data-level parallelism. The pipeline and instruction set of GAEA
are also optimized for the next generation wireless communication
systems. The evaluation and implementation results show that GAEA
architecture is suitable for the next generation wireless communication
systems.

#e1#

#s1#
#abstract1593.txt#


Increasing yield is important, especially for nano-scale technologies.
Also, pipelines are an important aspect of many SoC architectures. In
this paper we present new approaches to improve the yield and
yield/area of pipeline architectures by using (1) an appropriate number
of redundant copies for each module, and (2) sufficient steering logic
resources. We present an optimal algorithm of time complexity O(n/sup
3/) that adds redundant modules to an n-stage pipeline so as to
maximize yield. Experimental results indicate that for parameter values
of interests, this algorithm also improves the yield/area of the
pipeline, especially when the yield for some modules is low.

#e1#

#s1#
#abstract1594.txt#


Starting Electronic System Level (ESL) design flows with executable
High-Level Models (HLMs) has the potential to sustainability improve
productivity. However, writing good HLMs for complex systems is still a
challenging task. In the context of network controller design, modeling
complexity has two major sources: (1) the functionality to handle a
single connection, and (2) the number of connections to be handled in
parallel. In this paper, we will propose an efficient actor-oriented
modeling approach for complex systems by (1) integrating hierarchical
FSMs into dynamic dataflow models, and (2) providing new channel types
to allow concurrent processing of multiple connections. We will show
the applicability of our proposed modeling approach to real-world
system designs by presenting results from modeling and simulating a
network controller for the Parallel Sysplex architecture used in IBM
System z mainframes.

#e1#

#s1#
#abstract1595.txt#


We present a set of modeling constructs accompanied by a high
performance simulation kernel for accuracy adaptive transaction level
models. In contrast to traditional, fixed accuracy TLMs, accuracy of
adaptive TLMs can be changed during simulation to the level which is
most suitable for a given use case and scenario. Ad-hoc development of
adaptive models can result in complex models, and the implementation
detail of adaptivity mechanisms can obscure the actual logic of a
model. To simplify and enable systematic development of adaptive
models, we have identified several mechanisms which are applicable to a
wide variety of models. The proposed constructs relieve the modeler
from low level implementation details of those mechanisms. We have
developed an efficient, light-weight simulation kernel optimized for
the proposed constructs, which enables parallel simulation of large
models on widely available, low-cost multi-core simulation hosts. The
modeling constructs and the kernel have been evaluated using industrial
benchmark applications.

#e1#

#s1#
#abstract1596.txt#


Facing the requirements of next generation applications, current
approaches of embedded systems design will soon hit the limit where
they may no longer perform efficiently. The unpredictable nature and
diverse processing behavior of future applications requires to
transgress the barrier of tailor-made, application-/domain-specific
embedded system designs. As a consequence, next generation
architectures for embedded systems have to react much more flexible to
unforeseeable run-time scenarios. In this paper we present our
innovative processor architecture concept KAHRISMA (KArlsruhe's
Hypermorphic Reconfigurable-Instruction-Set Multi-grained-Array). It
tightly integrates coarse- and fine-grained run-time reconfigurable
fabrics that can incorporate to realize hardware acceleration for
computationally complex algorithms. Furthermore, the fabrics can be
combined to realize different Instruction Set Architectures that may
execute in parallel. With the help of an encrypted H.264 en-/decoding
case study we demonstrate that our novel KAHRISMA architecture will
deliver the required flexibility to design future-proof embedded
systems that are not limited to a certain computational domain.

#e1#

#s1#
#abstract1597.txt#


Modern automotive and aerospace embedded applications require very
high-performance simulations that are able to produce new values every
microsecond. Simulations must now rely on scalable performance of
multi-core systems rather than faster clock frequencies. Novel
parallelization techniques are needed to satisfy the industrial
simulation demands that are essential for the development of
safety-critical systems. Simulink formalism is the industrial de facto
standard, but current state-of-the-art simulation and code generation
techniques fail to fully exploit the parallelism in modern multi-core
systems. However, closed-loop and dynamic system simulations are very
difficult to parallelize because of the loop-carried dependencies. In
this paper we introduce a novel skewed pipelining technique that
overcomes these difficulties and allows loop-carried Simulink
applications to be executed concurrently in multi-core systems. By
delaying the forwarding of values for a few iterations, we can break
some data dependencies and coarsen the granularity of programs. This
improves the concurrency and reduces the high cost of inter-processor
communication. Implementation studies to demonstrate the viability of
our method on a commodity multi-core system with 2, 3, and 4 processors
show a 1.72, 2.38, and 3.33 fold speedup over uniprocessor execution.

#e1#

#s1#
#abstract1598.txt#


In conventional static implementations for correlated streaming
applications, computing resources may be in-efficiently utilized since
multiple stream processors may supply their sub-results at asynchronous
rates for result correlation or synchronization. To enhance the
resource utilization efficiency, we analyze multi-streaming models and
implement an adaptive architecture based on FPGA Partial
Reconfiguration (PR) technology. The adaptive system can intelligently
schedule and manage various processing modules during run-time.
Experimental results demonstrate up to 78.2% improvement in
throughput-per-unit-area on unbalanced processing of correlated
streams, as well as only 0.3% context switching overhead in the overall
processing time in the worst-case.

#e1#

#s1#
#abstract1599.txt#


We present a transactional datapath specification (T-spec) and the tool
(T-piper) to synthesize automatically an in-order pipelined
implementation from it. T-spec abstractly views a datapath as executing
one transaction at a time, computing next system states based on
current ones. From a T-spec, T-piper can synthesize a pipelined
implementation that preserves original transaction semantics, while
allowing simultaneous execution of multiple overlapped transactions
across pipeline stages. T-piper not only ensures the correctness of
pipelined executions, but can also employ forwarding and speculation to
minimize performance loss due to data dependencies. Design case studies
on RISC and CISC processor pipeline development are reported.

#e1#

#s1#
#abstract1600.txt#


Dynamic contrast-enhanced (DCE) magnetic resonance imaging (MRI) of the
prostate gland when evaluated along with T2-weighted images,
diffusion-weighted images (DWI) and their corresponding apparent
diffusion coefficient (ADC) maps can yield valuable information in
patients with rising or elevated serum prostate-specific antigen (PSA)
levels. In some cases, patients present with multiple negative
trans-rectal ultrasound (TRUS) biopsies, often placing the patient into
a cycle of active surveillance. Recently, more patients are undergoing
TRIM for targeted biopsy of suspicious findings with a cancer yield of
~59% compared to 15% for second TRUS biopsy2 to solve this diagnostic
dilemma and plan treatment. Patients were imaged in two separate
sessions on a 1.5T magnet using a cardiac phased array parallel imaging
coil. Automated CAD software was used to identify areas of wash-out. If
a suspicious finding was identified on all sequences it was followed by
a second imaging session. Under MRI-guidance, cores were acquired from
each target region. In one case the microscopic diagnosis was prostatic
intraepithelial neoplasia (PIN), in the other it was invasive
adenocarcinoma. Patient 1 had two negative TRUS biopsies and a PSA
level of 9ng/mL. Patient 2 had a PSA of 7.2ng/mL. He underwent TRUS
biopsy which was negative for malignancy. He was able to go on to
treatment for his prostate carcinoma (PCa)4. MRI may have an important
role in a subset of patients with multiple negative TRUS biopsies and
elevated or rising PSA.

#e1#

#s1#
#abstract1601.txt#


We introduce a novel solid modeling framework taking advantage of the
architecture of parallel computing on modern graphics hardware. Solid
models in this framework are represented by an extension of the ray
representation - Layered Depth-Normal Images (LDNI), which inherits the
good properties of Boolean simplicity, localization and domain
decoupling. The defect of ray representation in computational intensity
has been overcome by the newly developed parallel algorithms running on
the graphics hardware equipped with Graphics Processing Unit (GPU). The
LDNI for a solid model whose boundary is represented by a closed
polygonal mesh can be generated efficiently with the help of hardware
accelerated sampling. The parallel algorithm for computing Boolean
operations on two LDNI solids runs well on modern graphics hardware. A
parallel algorithm is also introduced in this paper to convert LDNI
solids to sharp-feature preserved polygonal mesh surfaces, which can be
used in downstream applications (e.g., finite element analysis).
Different from those GPU-based techniques for rendering CSG-tree of
solid models Hable and Rossignac (2007, 2005) [1,2], we compute and
store the shape of objects in solid modeling completely on graphics
hardware. This greatly eliminates the communication bottleneck between
the graphics memory and the main memory. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1602.txt#


BST thin films with various dopants were grown by the sol-gel method on
platinized silicon and MgO substrates. Their dielectric properties were
investigated at low frequency (up to 1 MHz) on silicon with
parallel-plate capacitors and at high frequency (up to 15 GHz) with
interdigitated capacitors on MgO substrate. The results depend on the
nature of the dopant and show that Mg is a very good candidate to
reduce dielectric losses. On the other hand, K is a good candidate as
dopant of BST thin film to drastically increase the tunability.

#e1#

#s1#
#abstract1603.txt#


A new scaled radix-4

#e1#

#s1#
#abstract1604.txt#


In this paper, we describe the first implementation and performance of
a fast O(N/sup 3/logN) hierarchical backprojection algorithm for cone
beam CT with a circular trajectory,developed on a modern Graphics
Processing Unit (GPU). The resulting tomographic backprojection system
for 3D cone beam geometry combines speedup through algorithmic
improvements provided by the hierarchical backprojection algorithm with
speedup from a massively parallel hardware accelerator. For data
parameters typical in diagnostic CT and using a mid-range GPU card, we
report reconstruction speeds of up to 360 frames per second, and
relative speedup of almost 6x compared to conventional backprojection
on the same hardware. The significance of these results is twofold.
First, they demonstrate that the reduction in operation counts
demonstrated previously for the FHBP algorithm can be translated to a
comparable run-time improvement in a massively parallel hardware
implementation, while preserving stringent diagnostic image quality.
Second, the dramatic speedup and throughput numbers achieved indicate
the feasibility of systems based on this technology, which achieve
real-time 3D reconstruction for state-of-the art diagnostic CT scanners
with small footprint, high-reliability, and affordable cost.

#e1#

#s1#
#abstract1624.txt#


Multi-cores are more and more popular recently and have being altered
the course of computing. Traditional XPath query evaluation algorithms
cannot take full advantages of multi-cores, and it is not
straightforward to adapt such algorithms on multi-cores. In this paper,
we propose an efficient parallel PathStack algorithm, named
P-PathStack, for processing XML twig queries. The algorithm first
efficiently partitions input element lists into multiple buckets, and
then processes data in each bucket in parallel. With efficient
partitioning method, our proposed algorithm can avoid many useless
elements and achieve very good speedup ratio. We have implemented the
algorithm and experimental results show that it achieves high
performance and speedup ratio.

#e1#

#s1#
#abstract1625.txt#


Generation of low-order harmonics (third and fifth) of the fundamental
radiation of a Q-switched Nd:YAG laser (1064 nm, pulse 15 ns) was
observed in a CaF/sub 2/ laser ablation plume. The ablation process is
triggered by a second Q-switched Nd:YAG laser operating at 532 or 266
nm. In the scheme employed, the fundamental laser beam propagates
parallel to the target surface at controllable distance and temporal
delay, allowing to the probing of different regions of the freely
expanding plume. The intensity of the harmonics is shown to decrease
rapidly as the distance to the target is increased, and for each
distance, an optimum time delay between the ablating laser pulse and
the fundamental beam is found. /i In situ/ diagnosis of the plume by
optical emission spectroscopy and laser-induced fluorescence serves to
correlate the observed harmonic behavior with the temporally and
spatially resolved composition and velocity of flight of species in the
plume. It is concluded that harmonics are selectively generated by CaF
species through a two-photon resonantly enhanced sum-mixing process
exploiting the (B/sup 2/ Sigma /sup +/-X/sup 2/ Sigma /sup +/, Delta nu
=0) transition of the molecule in the region of 530 nm. In this work
polar molecules have been shown to be the dominating species for
harmonic generation in an ablation plume. Implications of these results
for the generation of high harmonics in strongly polar molecules which
can be aligned in the ablation plasma are discussed.

#e1#

#s1#
#abstract1626.txt#


We present an algorithm for solid organ registration of pre-segmented
data represented as tetrahedral meshes. Registration of the organ
surface is driven by force terms based on a distance field
representation of the source and reference shapes. Registration of
internal morphology is achieved using a non-linear elastic finite
element model. A key feature of the method is that the user does not
need to specify boundary conditions (surface point correspondences)
prior to the finite element analysis. Instead the boundary matches are
found as an integrated part of the analysis. The method is evaluated on
phantom data and prostate data obtained in vivo based on fiducial
marker accuracy and inverse consistency of transformations. The
parallel nature of the method allows an efficient implementation on a
GPU and as a result the method is very fast. All validation
registrations take less than 30 seconds to complete. The proposed
method has many potential uses in image guided radiotherapy (IGRT)
which relies on registration to account for organ deformation between
treatment sessions.

#e1#

#s1#
#abstract1629.txt#


We consider the problem of scheduling a set of n independent jobs on m
parallel machines, where each job can only be scheduled on a subset of
machines called its processing set. The machines are linearly ordered,
and the processing set of job j is given by two machine indexes a/sub
j/ and b/sub j/; i.e., job j can only be scheduled on machines a/sub
j/, a/sub j/ + 1,..., b/sub j/. Two distinct processing sets are either
nested or disjoint. Preemption is not allowed. Our goal is to minimize
the makespan. It is known that the problem is strongly NP-hard and that
there is a list-type algorithm with a worst-case bound of 2-1/m. In
this paper we give an improved algorithm with a worst-case bound of
7/4. For two and three machines, the algorithm gives a better
worst-case bound of 5/4 and 3/2, respectively. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract1649.txt#


Multiple Sclerosis (MS) is a progressive neurological disease affecting
myelin pathways in the brain. Multiple lesions in the white matter can
cause paralysis and severe motor disabilities of the affected patient.
To solve the issue of inconsistency and user-dependency in manual
lesion measurement of MRI, we have proposed a 3-D automated lesion
quantification algorithm to enable objective and efficient lesion
volume tracking. The computer-aided detection (CAD) of MS, written in
MATLAB, utilizes K-Nearest Neighbors (KNN) method to compute the
probability of lesions on a per-voxel basis. Despite the highly
optimized algorithm of imaging processing that is used in CAD
development, MS CAD integration and evaluation in clinical workflow is
technically challenging due to the requirement of high computation
rates and memory bandwidth in the recursive nature of the algorithm. In
this paper, we present the development and evaluation of using a
computing engine in the graphical processing unit (GPU) with MATLAB for
segmentation of MS lesions. The paper investigates the utilization of a
high-end GPU for parallel computing of KNN in the MATLAB environment to
improve algorithm performance. The integration is accomplished using
NVIDIA's CUDA developmental toolkit for MATLAB. The results of this
study will validate the practicality and effectiveness of the prototype
MS CAD in a clinical setting. The GPU method may allow MS CAD to
rapidly integrate in an electronic patient record or any
disease-centric health care system.

#e1#

#s1#
#abstract1650.txt#


Multiple Sclerosis (MS) is a common neurological disease affecting the
central nervous system characterized by pathologic changes including
demyelination and axonal injury. MR imaging has become the most
important tool to evaluate the disease progression of MS which is
characterized by the occurrence of white matter lesions. Currently,
radiologists evaluate and assess the multiple sclerosis lesions
manually by estimating the lesion volume and amount of lesions. This
process is extremely time-consuming and sensitive to intra- and
inter-observer variability. Therefore, there is a need for automatic
segmentation of the MS lesions followed by lesion quantification. We
have developed a fully automatic segmentation algorithm to identify the
MS lesions. The segmentation algorithm is accelerated by parallel
computing using Graphics Processing Units (GPU) for practical
implementation into a clinical environment. Subsequently, characterized
quantification of the lesions is performed. The quantification results,
which include lesion volume and amount of lesions, are stored in a
structured report together with the lesion location in the brain to
establish a standardized representation of the disease progression of
the patient. The development of this structured report in collaboration
with radiologists aims to facilitate outcome analysis and treatment
assessment of the disease and will be standardized based on DI

#e1#

#s1#
#abstract1651.txt#


We present preliminary results obtained using a time domain wave-based
reconstruction algorithm for an ultrasound transmission tomography
scanner with a circular geometry. While a comprehensive description of
this type of algorithm has already been given elsewhere, the focus of
this work is on some practical issues arising with this approach. In
fact, wave-based reconstruction methods suffer from two major drawbacks
which limit their application in a practical setting: convergence is
difficult to obtain and the computational cost is prohibitive. We
address the first problem by appropriate initialization using a
ray-based reconstruction. Then, the complexity of the method is reduced
by means of an efficient parallel implementation on graphical
processing units (GPU). We provide a mathematical derivation of the
wave-based method under consideration, describe some details of our
implementation and present simulation results obtained with a numerical
phantom designed for a breast cancer detection application. The source
code of our GPU implementation is freely available on the web at
www.usense.org.

#e1#

#s1#
#abstract1652.txt#


This paper proposes a parallel computing method of multi-thread
technology replacing the former serial-execution one according to the
analysis of the radar imaging, the processing flow of NCS algorithm and
the characteristic of multiprocessor and multi-core system, considering
of the hard demand of high speed in imaging of raw spaceborne SAR data.
A multi-thread program based on Linux and the pthread encapsulated is
also provided. Through the experiments on HP Proliant DL580, it is
proved that, compared with traditional serial NCS algorithm, the method
of multi-thread presented in this paper makes use of system resource
effectively and is good for parallel expansibility of imaging as well
as achieving nice resolution.

#e1#

#s1#
#abstract1653.txt#


This paper presents a bi-mode, high performance Discrete Wavelet
Transform-based image compression core for future spacecrafts and
micro-satellites. The hardware solution proposed here exploits an
optimized CCSDS IDC algorithm and consists of a lossless mode and a
lossly mode. The throughput of the core is improved 40 times using
parallel architecture and pipeline technique, and a data rate of about
12.5M pixles/s can be sustained at 50MHz for the lossly mode. The
experimental results indicate that this core has high performance at
coding efficiency, data rate and error containment for both of the two
modes.

#e1#

#s1#
#abstract1654.txt#


Low density parity check (LDPC) codes are a class of
forward-error-correction codes. They are among the best-known codes
capable of achieving low bit error rates (BER) approaching Shannon's
capacity limit. Recently, LDPC codes have been adopted by the European
Digital Video Broadcasting (DVB-S2) standard, and have also been
proposed for the emerging IEEE 802.16 fixed and mobile broadband
wireless-access standard. The consultative committee for space data
system (CCSDS) has also recommended using LDPC codes in the deep space
communications and near-earth communications. It is obvious that LDPC
codes will be widely used in wired and wireless communication, magnetic
recording, optical networking, DVB, and other fields in the near
future. Efficient hardware implementation of LDPC codes is of great
interest since LDPC codes are being considered for a wide range of
applications. This paper presents an efficient partially parallel
decoder architecture suited for quasi-cyclic (QC) LDPC codes using
Belief propagation algorithm for decoding. Algorithmic transformation
and architectural level optimization are incorporated to reduce the
critical path. First, analyze the check matrix of LDPC code, to find
out the relationship between the row weight and the column weight. And
then, the sharing level of the check node updating units (CNU) and the
variable node updating units (VNU) are determined according to the
relationship. After that, rearrange the CNU and the VNU, and divide
them into several smaller parts, with the help of some assistant logic
circuit, these smaller parts can be grouped into CNU during the check
node update processing and grouped into VNU during the variable node
update processing. These smaller parts are called node update kernel
units (NKU) and the assistant logic circuit are called node update
auxiliary unit (NAU). With NAUs' help, the two steps of iteration
operation are completed by NKUs, which brings in great hardware
resource reduction. Meanwhile, efficient techniques have been developed
to reduce the computation delay of the node processing units and to
minimize hardware overhead for parallel processing. This method may be
applied not only to regular LDPC codes, but also to the irregular ones.
Based on the proposed architectures, a (7493, 6096) irregular QC-LDPC
code decoder is described using verilog hardware design language and
implemented on Altera field programmable gate array (FPGA) StratixII
EP2S130. The implementation results show that over 20% of logic core
size can be saved than conventional partially parallel decoder
architectures without any performance degradation. If the decoding
clock is 100MHz, the proposed decoder can achieve a maximum (source
data) decoding throughput of 133 Mb/s at 18 iterations.

#e1#

#s1#
#abstract1655.txt#


User authentication using fingerprint information provides convenience
as well as strong security at the same time. However, serious problems
may cause if fingerprint information stored for user authentication is
used illegally by a different person since it cannot be changed freely
as a password due to a limited number of fingers. Recently, research in
fuzzy fingerprint vault system has been carried out actively to safely
protect fingerprint information in a fingerprint authentication system.
In this paper, we propose hardware architecture for a geometric hashing
based fuzzy fingerprint vault system. The proposed architecture
consists of the software module and hardware module. The hardware
module performs the matching for the transformed minutiae in the
enrollment and verification hash table. We also propose a hardware
architecture which parallel processing technique is applied for high
speed processing.

#e1#

#s1#
#abstract1656.txt#


Optical lithography, currently being used for 45-nm semiconductor
devices, is expected to be extended further towards the 32-nm and 22-nm
node. A further increase of lens NA will not be possible but
fortunately the shrink can be enabled with new resolution enhancement
methods like source mask optimization (SMO) and double patterning
techniques (DPT). These new applications lower the k1 dramatically and
require very tight overlay control and CD control to be successful. In
addition, overall cost per wafer needs to be lowered to make the
production of semiconductor devices acceptable. For this ultimate era
of optical lithography we have developed the next generation dual stage
NXT:1950i immersion platform. This system delivers wafer throughput of
175 wafers per hour together with an overlay of 2.5 nm. Several
extensions are offered enabling 200 wafers per hour and improved
imaging and on product overlay. The high productivity is achieved using
a dual wafer stage with planar motor that enables a high acceleration
and high scan speed. With the dual stage concept wafer metrology is
performed in parallel with the wafer exposure. The free moving planar
stage has reduced overhead during chuck exchange which also improves
litho tool productivity. In general, overlay contributors are coming
from the lithography system, the mask and the processing. Main
contributors for the scanner system are thermal wafer and stage
control, lens aberration control, stage positioning and alignment. The
back-bone of the NXT:1950i enhanced overlay performance is the novel
short beam fixed length encoder grid-plate positioning system. By
eliminating the variable length interferometer system used in the
previous generation scanners the sensitivity to thermal and flow
disturbances are largely reduced. The alignment accuracy and the
alignment sensitivity for process layers are improved with the SMASH
alignment sensor. A high number of alignment marker pairs can be used
without throughput loss, and furthermore the GridMapper functionality
which is using the inter-die and intra-die scanner capability can
reduce overlay errors coming from mask and process without productivity
impact. In this paper we will present the main design features and
discuss the system performance of the NXT:1950i system, focusing on the
improvements made in overlay and productivity. We will show data on
imaging, overlay, focus and productivity supporting the 3X-nm node and
we will discuss next improvement steps towards the 2X-nm node.

#e1#

#s1#
#abstract1657.txt#


Early detection, diagnosis, and suitable treatment are known to
significantly improve the chance of survival for breast cancer (BC)
patients. To date, the most cost effective method for screening and
early detection is mammography, which is also the tool that has
demonstrated its ability to reduce BC mortality. Tomosynthesis is an
emerging technology that offers an alternative to conventional
two-dimensional mammography. Tomosynthesis produces three-dimensional
(volumetric) images of the breast that may be superior to planar
imaging due to improved visualization. In this paper we examined the
effect of varying the number of projections (N) and total view angle
(VA) on the shift-and-add (SAA), back projection (BP) and filtered back
projection (FBP) image reconstruction response characterized by impulse
response (IR) simulations. IR data were generated by simulating the
projection images of a very thin wire, using various combinations of VA
and N. Results suggested that BP and FBP performed better for in-plane
performance than that of SAA. With bigger number of projection images,
the investigated reconstruction algorithms performed the best by
obtaining sharper in-focus IR with simulated parallel imaging
configurations.

#e1#

#s1#
#abstract1658.txt#


The purpose of this study was to develop and implement an accurate and
computationally efficient method for determination of the mesh-domain
system matrix including attenuation compensation for Ordered Subsets
Expectation Maximization (OSEM) Single Photon Emission Computed
Tomography (SPECT). The mesh-domain system matrix elements were
estimated by first partitioning the object domain into strips parallel
to detector face and with width not exceeding the size of a detector
unit. This was followed by approximating the integration over the
strip/mesh-element union. This approximation is product of: (i) strip
width, (ii) intersection length of a ray central to strip with a mesh
element, and (iii) the response and expansion function evaluated at
midpoint of the intersection length. Reconstruction was performed using
OSEM without regularization and with exact knowledge of the attenuation
map. The method was evaluated using synthetic SPECT data generated
using SIMIND Monte Carlo simulation software. Comparative quantitative
and qualitative analysis included: bias, variance, standard deviation
and line-profiles within three different regions of interest. We found
that no more than two divisions per detector bin were needed for good
quality reconstructed images when using a high resolution mesh.

#e1#

#s1#
#abstract1659.txt#


Parallel magnetic resonance imaging achieves reduction in scan time by
collecting a partial set of signals using an array of receiving coils
each with a local sensitivity pattern. An image is then reconstructed
from the partial dataset using the additional information of coil
sensitivity. GRAPPA (generalized auto calibrating partially parallel
acquisitions) is one of the most successful reconstruction techniques
in which the missing k-space lines are interpolated from the acquired
data in the whole coil array using a convolution kernel estimated from
a fully sampled data patch in the center of k-space. The interpolation
kernel is usually small but fixed in size for all coils. Here, we show
that a variable kernel with a size dependent on the coil sensitivity
can lead to better image quality. The kernel size is estimated from the
ratio of the coil sensitivities obtained from a reference scan or from
the same dataset. Conventional GRAPPA kernel estimation and image
reconstruction is modified to employ the variable-size kernel for
improved reconstruction. The new technique shows improved image quality
compared to GRAPPA.

#e1#

#s1#
#abstract1660.txt#


Ring artifacts often appear in flat-detector CT because of imperfect or
defect detector elements or calibration. In high-spatial resolution CT
images reducing such artifacts becomes a necessity. In this paper, we
used the post-processing ring correction in polar coordinates (RCP) to
eliminate the ring artifacts. The median filter is applied to the
uncorrected images in polar coordinates and ring artifacts are
extracted from the original images. The algorithm has a very high
computational cost due to the time-expensive median filtering and
coordinate transformation on CPUs. Graphics processing units (GPUs)ca n
be seen as parallel co-processors with high computational power. All
steps of the RCP algorithm were implemented with CUDA (Compute Unified
Device Architecture, NVIDIA). We introduced a new GPU-based branchless
vectorized median (BVM) filter. This algorithm is based on minmax
sorting and keeps track of a sorted array from which values are deleted
and to which new values are inserted. For comparison purpose a modified
pivot median filter on GPUs was presented, which compares a pivot
element to all other values and recursively finds the median element.
We evaluated the performance of the RCP method using 512 slices, each
slice consisted of 512 * 512 pixels. This post-processing method
efficiently reduces ring artifacts in the reconstructed images and
improves image quality. Our CUDA-based RCP is up to 13.6 times faster
than the optimized CPU-based (single core)r outine. Comparing our two
GPU-based median filters showed a performance benefit by roughly 60%
when switching from Pivot to BVM code. The main reason is that the BVM
algorithm is branchless and makes use of data-level parallelism. The
BVM method is better suited to the model of modern graphics processing.
A multi-GPU solution showed that the performance scaled nearly linearly.

#e1#

#s1#
#abstract1661.txt#


The simulation of imaging systems using Monte Carlo x-ray transport
codes is a computationally intensive task. Typically, many days of
computation are required to simulate a radiographic projection image
and, as a consequence, the simulation of the hundreds of projections
needed to perform a tomographic reconstruction may require an
unaffordable amount of computing time. To speed up x-ray transport
simulations, a MC code that can be executed in a graphics processing
unit (GPU) was developed using the CUDATM programming model, an
extension to the C language for the execution of general-purpose
computations on NVIDIA's GPUs. The code implements the accurate photon
interaction models from PENELOPE and takes full advantage of the GPU
massively parallel architecture by simulating hundreds of particle
tracks simultaneously. In this work we describe a new version of this
code adapted to the simulation of computed tomography (CT) scans, and
allowing the execution in parallel in multiple GPUs. An example
simulation of a cardiac CT using a detailed voxelized anthropomorphic
phantom is presented. A comparison of the simulation computational
performance in one or multiple GPUs and in a CPU (Central Processing
Unit), and a benchmark with a standard PENELOPE code, are provided.
This study shows that low-cost GPU clusters are a good alternative to
CPU clusters for Monte Carlo simulation of x-ray transport.

#e1#

#s1#
#abstract1662.txt#


Novel geometrical designs of computed tomography (CT) scanners in
combination with novel image reconstruction algorithms promise to
reduce ionizing radiation exposure to the patient in CT scans. While
the sampling density of the Field Of View (FOV) is retained, the image
quality can even be increased in contrast to conventional CT scanners.
In this study, we present first images obtained with a novel CT scanner
that we developed in our working group. In this open CT system with
irradiation within a fan beam, parallel Radon data are directly
obtained for image reconstruction using the OPED (Orthogonal Polynomial
Expansion on the Disk) algorithm. This algorithm uses Radon data
directly, i.e., without any further data processing such as rebinning
and interpolation. We experimentally test theoretical predictions for
this system by quantifying image quality parameters in comparison with
corresponding parameters that are derived from the images of a
conventional scanner of the 3rd generation. The modulation transfer
function (MTF) and noise power spectrum (NPS) are determined using a
test phantom. The novel CT system quantitatively shows the same noise
property as the conventional scanner. The resolution that is reached in
the center of a reconstructed image is nearly identical for both
scanner types. But we found that the resolution that is achieved in the
novel CT system does not depend on the image position while the MTF of
the conventional scanner decreases for radially outer regions of the
image.

#e1#

#s1#
#abstract1663.txt#


An angular parameterization of parallel Radon projections referred to
in this paper as psi -parameterization is discussed in relevance to the
efficiency of reconstruction from fan data. The fact that the psi
-parameterization coincides with the equiangular fan beam
parameterization allows us to develop a simple and efficient approach
useful for the reconstruction from fan data. Within this approach
parallel projections are approximated by groups of semi-parallel rays.
The reconstruction is carried out directly, i.e. without any
modification of original data, at the speed which is comparable or even
higher than that of the parallel Filtered Back Projection (FBP)
algorithm.

#e1#

#s1#
#abstract1664.txt#


Digital breast tomosynthesis is a new technique to improve the early
detection of breast cancer by providing three-dimensional
reconstruction volume of the object with limited-angle projection
images. This paper investigated the image reconstruction with a
standard biopsy training breast phantom using a novel multi-beam X-ray
sources breast tomosynthesis system. Carbon nanotube technology based
X-ray tubes were lined up along a parallel-imaging geometry to decrease
the motion blur. Five representative reconstruction algorithms,
including back projection (BP), filtered back projection (FBP), matrix
inversion tomosynthesis (MITS), maximum likelihood expectation
maximization (MLEM) and simultaneous algebraic reconstruction technique
(SART), were investigated to evaluate the image reconstruction of the
tomosynthesis system. Reconstructed images of the masses and
micro-calcification clusters embedded in the phantom were studied. The
evaluated multi-beam X-ray breast tomosynthesis system is able to
generate three-dimensional information of the breast phantom with
clearly-identified regions of the masses and calcifications. Future
study will be done soon to further improve the imaging parameters'
measurement and reconstruction.

#e1#

#s1#
#abstract1665.txt#


Introduction: Titanium implants can be regarded as the current gold
standard for restoration of sound transmission in the middle ear
following destruction of the ossicular chain by chronic inflammation.
Many efforts have been made to improve prosthesis design, while less
attention had been given to the role of the interface. We present a
study on chemical nanocoating on microstructured titanium contact
surface with bioactive protein. Materials and Methods: Titanium samples
of 5 mm diameter and 0,25 mm thickness were structured by means of a
Ti:Sapphire femtosecond laser operating at 970 nm with parallel lines
of 5 mu m depth, 5 mu m width and 10 mu m inter-groove distance. In
addition, various nanolayers were applied to titanium samples by
aminosilanization, to which Star-Polyethylene glycole (Star-PEG)
molecules plus biomarkers (e.g. RGD peptide sequence) were linked.
Results: Chondrocytes could be cultured on microstructured surfaces
without reduced rate of vital / dead cells compared to native surfaces.
Chondrocytes also showed contact guidance by growing along ridges
particularly on 5 mu m lines. On nanocoated titanium samples, first
results showed a strong effect of Star-PEG suppressing unspecific
protein absorption, while RGD peptide sequence did not promote
chondrocyte cell growth. Discussion: According to these results, the
idea of promoting cell growth on titanium prosthesis contact surfaces
compared to non-contact surfaces (e.g. prosthesis shaft) by nanocoating
is practicable. However, relative selectivity induced by
microstructures for growth of chondrocytes compared to fibrocytes is
subject to further evaluation.

#e1#

#s1#
#abstract1666.txt#


Quantification of vessel wall thickness is important in longitudinal
monitoring of atherosclerosis. Black-blood MRI has been useful in
measuring vessel wall thickness. Studies using two-dimensional (2D)
imaging protocols measured wall thickness by matching the arterial wall
and lumen boundaries on an acquisition plane. If the acquisition plane
is oblique to the artery, the wall thickness would be overestimated by
a factor that is dependent on the obliqueness angle. This problem can
be understood as a three-dimensional (3D) surface mismatch problem, and
we evaluated the effect of this problem by comparing the thickness
measurements obtained using a 2D contour matching method and a 3D
surface matching method. In addition to the surface mismatch problem,
two other parameters may affect the wall thickness estimation:
reslicing angle and slice thickness. We measured the wall thickness
using images resliced perpendicular to the centerline of the vessel and
quantified the difference between the thickness measurements obtained
from parallel and centerline-based resliced images. Images obtained
from a 2D MRI protocol typically have a slice thickness of 2 mm, while
the 3D MRI technique applied in this study produced images with
sub-millimeter isotropic voxel size. To investigate the effect of slice
thickness, we simulated 2 mm-thick images by averaging the 3D
black-blood image. Our results show that the wall thickness measured
from 2 mm-thick images was overestimated, especially in the carotid
artery, which is associated with a larger obliqueness angle. This
result underscores the advantage of the 3D isotropic acquisition
technique in wall thickness measurement, especially in more tortuous
vessels.

#e1#

#s1#
#abstract1667.txt#


Fast Fourier Transform (FFT) is an important algorithm in many digital
signal processing applications, and it often requires parallel
implementation for high throughput. In this paper, we first present the
SmartCell coarse-grained reconfigurable architecture targeted for
stream processing. A SmartCell prototype integrates 64 processing
elements, configurable interconnections, and dedicated instruction and
data memories into a single chip, which is able to provide high
performance parallel processing while maintaining post-fabrication
flexibility. Subsequently, we present a parallel FFT architecture
targeted for multi-core platforms computing systems. This algorithm
provides an optimized data flow pattern that reduces both communication
and configuration overheads. The proposed parallel FFT algorithm is
then mapped onto the SmartCell prototype device. Results show that the
parallel FFT implementation on SmartCell is about 14.9 and 2.7 times
faster than network-on-chip (NoC) and MorphoSys implementations,
respectively. SmartCell also achieves the energy efficiency gains of
2.1 and 28.9 when compared with FPGA and DSP implementations.

#e1#

#s1#
#abstract1668.txt#


One Super Hi-Vision (SHV) 4 k * 4 k @ 60 fps fractional motion
estimation (FME) engine is proposed in our paper. Firstly, two
complexity reduction schemes are proposed in the algorithm level. By
analyzing the integer motion cost of sub blocks in each inter mode, the
mode reduction based mode pre-filtering scheme can achieve 48% clock
cycle saving compared with previous algorithm. By further check the
motion cost of search points around best integer candidate, the motion
cost oriented directional one-pass scheme can provide 50% clock cycle
saving and 36% reduction in the number of processing units (PU).
Secondly, in the hardware level, two parallel improved schemes namely
16-Pel processing and MB-parallel scheme are given out in our paper,
which reduces design effort to only 145 MHz for SHV FME processing.
Also, quarter sub-sampling is adopted in our design and 75% hardware
cost is reduced for each PU. Thirdly, one unified pixel block loading
scheme is proposed. About 28.67% to 86.39% pixels are reused and the
related memory access is saved. Furthermore, we also give out one
parity pixel organization scheme to solve memory access conflict of
MB-parallel scheme. By using TSMC 0.18 mu m technology in worst work
conditions (1.62 V, 125 degrees C), our FME engine can achieve
real-time processing for SHV 4 k * 4 k @ 60 fps with 412 k gates
hardware.

#e1#

#s1#
#abstract1670.txt#


In this paper we introduce new algorithm implementations of a new
parametric image processing framework that will accurately process
images and speed up computation for addition, subtraction, and
multiplication. Its potential applications include computer graphics,
digital signal processing and other multimedia applications. This
Parameterized Digital Electronic Arithmetic (PDEA) model replaces
linear operations with non-linear ones. The implementation of a
parameterized model is presented. We also present the design of
arithmetic circuits including parallel counters, adders and multipliers
based in two high performance threshold logic gate implementations that
we have developed. We will also explore new microprocessor
architectures to take advantage of arithmetic. The experiments executed
have shown that the algorithm provides faster and better enhancements
from those described in the literature. The FPGA chips used is Spartan
3E from Xilinix. The critical length in the circuit implemented on the
FPGA had the minimum period for the proposed subsystem is 10.209 ns
(maximum frequency 97.957 MHz). Maximum power consumed is 2.4 mW using
32 nm process and we used parallelism and reuse of the Hardware
components to accomplish and speed up the process. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract1674.txt#


We propose a scheme for parallel spatially multimode quantum memory for
light. The scheme is based on a counterpropagating quantum signal wave
and a strong classical reference wave as in a classical volume hologram
and therefore can be called a /i quantum volume hologram/. The medium
for the hologram consists of a spatially extended ensemble of atoms
placed in a magnetic field. The write-in and readout of this quantum
hologram is as simple as that of its classical counterpart and consists
of a single-pass illumination. In addition, we show that the present
scheme for a quantum hologram is less sensitive to diffraction and
therefore is capable of achieving a higher density of storage of
spatial modes as compared to previous proposals. We present a
feasibility study and show that experimental implementation is possible
with available cold atomic samples. A quantum hologram capable of
storing entangled images can become an important ingredient in quantum
information processing and quantum imaging.

#e1#

#s1#
#abstract1675.txt#


This Rapid Communication presents a method of beam-divergence
deconvolution for diffractive imaging. First, the detected diffraction
intensity is formulated as a convolution between the diffraction
intensity of parallel incident beams and the divergence of an incident
beam. It is shown numerically that the convolution causes the
reconstructed image to shrink and become blurred. Next, the algorithm
of deconvolution used in the iterative Fourier phase retrieval method
is applied to the convoluted diffraction intensity deteriorated by
quantum noise. Numerical simulations show that the proposed algorithm
recovers the deconvoluted diffraction intensity and improves the
reconstructed image. Finally, the algorithm is applied to an
electron-beam experiment to reconstruct a multiwall carbon nanotube.
The results verified that the algorithm reduces the influence of beam
divergence.

#e1#

#s1#
#abstract1676.txt#


Using /i in situ/ electron diffraction we study the orientation of
mass-selected iron nanoparticles upon deposition onto single
crystalline W(110) at room temperature. It is found that particles with
a diameter below about 4 nm and a kinetic energy #or=0.1 electron volt
per atom spontaneously align with respect to the substrate. Larger
particles preferentially rest with their (001) and (110) facets
parallel to the surface, but do not show further alignment. The data
may hint at thermally activated dislocation motions upon the impact on
the substrate which are responsible for the observed orientation below
4 nm. By this uniformly oriented monodisperse nanostructures can be
prepared on single-crystalline substrates.

#e1#

#s1#
#abstract1677.txt#


Liver segmentation remains a difficult problem in medical images
processing, especially when accuracy and speed are both seriously
considered. Graph Cuts is a powerful segmentation tool through which
the optimal results are got by considering both region and boundary
information in images. However, the traditional Graph Cuts algorithms
are always computationally expensive and inappropriate to be applied to
real clinical circumstance. Recently, the GPU (Graphics Processor Unit)
had evolved to be a cheap and superpower general purpose computing
instrument, especially when NVIDIA released its revolutionary CUDA
(Compute Unified Device Architecture). In this paper, we introduce a
novel method to segment 3D liver images with GPU, using the
Push-Relable style 3D Graph Cuts implementation. Some modifications
such as 3D storage structures are also introduced which make our
implement well fit to the GPU parallel computing capabilities.
Experiments have been executed on human liver CT data and these
experiments show that our method can obtains results in much less time
compared to the implement with CPU.

#e1#

#s1#
#abstract1678.txt#


Image labeling and parcellation are critical tasks for the assessment
of volumetric and morphometric features in medical imaging data. The
process of image labeling is inherently error prone as images are
corrupted by noise and artifact. Even expert interpretations are
subject to subjectivity and the precision of the individual raters.
Hence, all labels must be considered imperfect with some degree of
inherent variability. One may seek multiple independent assessments to
both reduce this variability as well as quantify the degree of
uncertainty. Existing techniques exploit maximum a posteriori
statistics to combine data from multiple raters. A current limitation
with these approaches is that they require each rater to generate a
complete dataset, which is often impossible given both human foibles
and the typical turnover rate of raters in a research or clinical
environment. Herein, we propose a robust set of extensions that allow
for missing data, account for repeated label sets, and utilize
training/catch trial data. With these extensions, numerous raters can
label small, overlapping portions of a large dataset, and rater
heterogeneity can be robustly controlled while simultaneously
estimating a single, reliable label set and characterizing uncertainty.
The proposed approach enables parallel processing of labeling tasks and
reduces the otherwise detrimental impact of rater unavailability.

#e1#

#s1#
#abstract1680.txt#


Abstract: As the most accurate model for simulating light propagation
in heterogeneous tissues, Monte Carlo (MC) method has been widely used
in the field of optical molecular imaging. However, MC method is
timeconsuming due to the calculations of a large number of photons
propagation in tissues. The structural complexity of the heterogeneous
tissues further increases the computational time. In this paper we
present a parallel implementation for MC simulation of light
propagation in heterogeneous tissues whose surfaces are constructed by
different number of triangle meshes. On the basis of graphics
processing units (GPU), the code is implemented with compute unified
device architecture (CUDA) platform and optimized to reduce the access
latency as much as possible by making full use of the constant memory
and texture memory on GPU. We test the implementation in the
homogeneous and heterogeneous mouse models with a NVIDIA GTX 260 card
and a 2.40 GHz Intel Xeon CPU. The experimental results demonstrate the
feasibility and efficiency of the parallel MC simulation on GPU.

#e1#

#s1#
#abstract1681.txt#


We report a simple implementation to acquire spectral domain
polarization-sensitive optical coherence tomography (PSOCT) using a
single camera. By combining a dual-delay assembly in the reference arm
and offset B-scan in the sample arm, the orthogonal vertical- and
horizontal-polarized images were acquired in parallel and spatially
separated by a fixed distance in the full range image space. The two
orthogonal polarization images were recombined to calculate the
intensity, retardance and fast-axis images. This system was easy to
implement and capable of acquiring highspeed in vivo 3D
polarization-sensitive OCT images.

#e1#

#s1#
#abstract1682.txt#


With fast development of GPU hardware and software, using CPUs to
accelerate non-graphics CPU applications is becoming inevitable trend.
CPUs are good at performing ALU-intensive computation and feature high
peak performance; however, how to harness CPUs' powerful computing
capacity to accelerate the applications in the field of scientific
computing still remains a big challenge. In this paper, we implement
the whole application Mgrid taken from Spec2000 benchmarks on an AMD
GPU and propose several optimization strategies for stencil
computations in the naive GPU code. We first improve thread utilization
through using vector types and multiple output streams mechanism
provided by the Brook+ programming language. By tuning thread
granularity, we try to hit the right balance between locality within
each thread and parallelism among threads. Then, we reorganize the
stream layout by transforming the 3D data stream into the 2D stream in
the block manner. Through stream reorganization, more data locality in
the cache is exploited. Further, we propose branch elimination to
convert control dependence to data dependence, catering to GPUs'
powerful ALU-intensive processing capability. Finally, we redistribute
computations between CPU and GPU to make more advisable computing
resources usage considering different problem sizes. We demonstrate the
effectiveness of our proposed optimization strategies on an AMD Radeon
HD4870 GPU using the Brook+ programming language. Using a
double-precision floating-point implementation, the experimental
results show that the optimized GPU version of Mgrid gains 2.38x
speedup compared to the naive GPU code and obtains as high as 15.06x
speedup versus the CPU implementation run on an Intel Xeon E5405 CPU.

#e1#

#s1#
#abstract1683.txt#


Efficient transaction nesting is one of the ongoing challenges for
hardware transactional memory. To increase efficiency of closed
nesting, this paper proposes a conditional partial rollback (CPR)
scheme which supports conditional partial rollback without increasing
hardware complexities significantly. In stead of rolling back to the
outermost transaction as in commonly-used flattening model, the CPR
scheme just rolls back to the conflicted transaction itself or one of
its outer-level transactions if given conditions are satisfied. By
recording access status of each nested transaction, the scheme uses one
global data set for all of the nested transactions rather than
independent data set for each nested transaction. Hardware
transactional memory architecture with the support of CPR scheme is
also proposed based on multi-core processor and current cache coherence
mechanism. The system is implemented by simulation, and evaluated using
seven benchmark applications. Evaluation results show that the CPR
scheme achieves better performance and scalability than the flattening
model which is commonly-used in hardware transactional memory.

#e1#

#s1#
#abstract1684.txt#


The objective of Eurogene is to collect a critical mass of educational
content in the field of human genetics in nine European languages and
to build a platform that will support the retrieval, sharing and
navigation over the learning content. The Eurogene platform is already
operational and is being used by the genetics community. In this paper,
a part of the Eurogene platform related to the retrieval and machine
translation of domain specific content is described. Our contribution
lies in an approach for domain-specific adaption of cross-language
information retrieval (CLIR) and machine translation (MT). The CLIR
system is based on a multilingual domain ontology which is also used as
a synchronization component between CLIR and MT. The MT system is
adapted to the target domain using the terminology represented in the
ontology and using statistical training performed on a collection of
parallel texts. In the statistical training phase, new translations of
a term can be discovered and used for ontology updating. The paper is
organized as follows. First, we describe the motivation for our
approach and the multilingual domain ontology. Later, the CLIR and MT
components and their domain adaption and synchronization are discussed.

#e1#

#s1#
#abstract1685.txt#


We address the problem of Transliteration Equivalence, i.e. determining
whether a pair of words in two different languages (e.g. Auden) are
name transliterations or not. This problem is at the heart of Mining
Name Transliterations (MINT) from various sources of multilingual text
data including parallel, comparable, and non-comparable corpora and
multilingual news streams. MINT is useful in several cross-language
tasks including Cross-Language Information Retrieval (CLIR), Machine
Translation (MT), and Cross-Language Named Entity Retrieval. We propose
a novel approach to Transliteration Equivalence using language-neutral
representations of names. The key idea is to consider name
transliterations in two languages as two views of the same semantic
object and compute a low-dimensional common feature space using
Canonical Correlation Analysis (CCA). Similarity of the names in the
common feature space forms the basis for classifying a pair of names as
transliterations. We show that our approach outperforms
state-of-the-art baselines in the CLIR task for Hindi-English (3
collections) and Tamil-English (2 collections).

#e1#

#s1#
#abstract1686.txt#


Background: DNA signatures are distinct short nucleotide sequences that
provide valuable information that is used for various purposes, such as
the design of Polymerase Chain Reaction primers and microarray
experiments. Biologists usually use a discovery algorithm to find
unique signatures from DNA databases, and then apply the signatures to
microarray experiments. Such discovery algorithms require to set some
input factors, such as signature length l and mismatch tolerance d,
which affect the discovery results. However, suggestions about how to
select proper factor values are rare, especially when an unfamiliar DNA
database is used. In most cases, biologists typically select factor
values based on experience, or even by guessing. If the discovered
result is unsatisfactory, biologists change the input factors of the
algorithm to obtain a new result. This process is repeated until a
proper result is obtained. Implicit signatures under the discovery
condition (l, d) are defined as the signatures of length #or= l with
mismatch tolerance #or= d. A discovery algorithm that could discover
all implicit signatures, such that those that meet the requirements
concerning the results, would be more helpful than one that depends on
trial and error. However, existing discovery algorithms do not address
the need to discover all implicit signatures. Results: This work
proposes two discovery algorithms - the consecutive multiple discovery
(CMD) algorithm and the parallel and incremental signature discovery
(PISD) algorithm. The PISD algorithm is designed for efficiently
discovering signatures under a certain discovery condition. The
algorithm finds new results by using previously discovered results as
candidates, rather than by using the whole database. The PISD algorithm
further increases discovery efficiency by applying parallel computing.
The CMD algorithm is designed to discover implicit signatures
efficiently. It uses the PISD algorithm as a kernel routine to discover
implicit signatures efficiently under every feasible discovery
condition. Conclusions: The proposed algorithms discover implicit
signatures efficiently. The presented CMD algorithm has up to 97% less
execution time than typical sequential discovery algorithms in the
discovery of implicit signatures in experiments, when eight processing
cores are used.

#e1#

#s1#
#abstract1687.txt#


The following topics are dealt with: image processing; image analysis;
education; information shifting; software development; wireless
network; information system; biosensor; distributed computing;
intelligent system; emotion recognition; and parallel computing.
UT INSPEC:11271403

#e1#

#s1#
#abstract1688.txt#


The computational complexity of scheduling jobs with released dates on
an unbounded batch processing machine to minimize total completion time
and on parallel unbounded batch processing machines to minimize total
weighted completion time remains open. In this note we show that the
first problem is NP-hard with respect to id-encoding, and the second
one is strongly NP-hard. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1689.txt#


Biomedical sensors combining microfluidic and electronics capabilities
require defect avoidance in both the electronic processing circuits and
microfluidic areas. Microfluidic sensors involve sealed channels
through which sample fluids containing biomedical materials flow.
Inserting microchannels between capacitive plates enable the detection
of biomaterials by the changes in capacitance. However, faults occur
when foreign particles, or fluid bubbles get lodged in the paths
blocking a channel, thereby affecting the measured C. To achieve fault
tolerance we investigate a Cathedral Chamber design, with pillars
supporting the roof at regular intervals. This prevents single
blockages from stopping fluid flow through the system in a channel, as
there are many paths. We discuss the potential causes and effects of
such blockages. Monte Carlo simulations show that the Cathedral Chamber
design significantly increases lifetime of the system, an average of 6
times more particles are required before full blockage occurs compared
to an array of parallel channels. Fluid flow modeling shows parallel
channels show rapid rise of pressure with the number of blockages while
the Cathedral chamber shows much slower rise, which reaches a plateau
pressure until it is blocked. The impact of defects on the capacitive
measurement is also discussed. Finally, an interesting application, one
that uses patches of single chain Fragment variables (scFv's), the
active part of antibodies, is also discussed.

#e1#

#s1#
#abstract1690.txt#


Calculating decision rules is a very important process. There are a lot
of solutions for computing decision rules, but the algorithms of
computing all set of rules are time consuming. We propose a recursive
version of the well known apriori algorithm, designed for parallel
processing. We present here, how to decompose the problem of
calculating decision rules, so that the parallel calculations are
efficient.

#e1#

#s1#
#abstract1691.txt#


Joins between data sources are an essential ingredient of multidomain
queries, as they exploit connection patterns defined between service
marts or between service interfaces. This chapter moves from the
definition of a query language over service interfaces, sketching how
queries can be directly expressed over service marts and how these can
be translated over service interfaces. The fundamental operation
discussed in this chapter is the binary join between two sources, which
is influenced by the type (search vs. exact) of services and by the
management (parallel vs. sequential) of service calls. Then, this
chapter presents an optimization framework for queries over several
service interfaces, which considers several cost metrics for mapping
queries into query plans, consisting of specific operations over
services, and includes a branch and bound approach to the exploration
of the combinatorial search space of all possible query plans.

#e1#

#s1#
#abstract1692.txt#


Light scattering spectroscopy (LSS) and Fourier domain low coherence
interferometry (fLCI) are used in combination with the dual window
method (DW) to measure scattering features from a thick turbid sample.
By processing with the DW method, the trade off that hinders
spectroscopic OCT is avoided, thus yielding depth resolved spectra with
simultaneously high spatial and spectral resolution. The capabilities
of the method are demonstrated by analyzing a double layer phantom,
where the top layer contains polystyrene beads of diameter d = 4.00 mu
m, and the bottom layer contains beads of d = 6.98 mu m. A white light
parallel frequency domain OCT system is used to image the sample. The
results show that scattering structure can be assessed accurately and
precisely throughout the whole OCT image using LSS and fLCI.

#e1#

#s1#
#abstract1709.txt#


We study machine scheduling problems in which the jobs belong to
different job classes and they need to be delivered to customers after
processing. A setup time is required for a job if it is the first job
to be processed on a machine or its processing on a machine follows a
job that belongs to another class. Processed jobs are delivered in
batches to their respective customers. The batch size is limited by the
capacity of the delivery vehicles and each shipment incurs a transport
cost and takes a fixed amount of time. The objective is to minimize the
weighted sum of the last arrival time of jobs to customers and the
delivery (transportation) cost. For the problem of processing jobs on a
single machine and delivering them to multiple customers, we develop a
dynamic programming algorithm to solve the problem optimally. For the
problem of processing jobs on parallel machines and delivering them to
a single customer, we propose a heuristic and analyze its performance
bound. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1710.txt#


In practice, machine schedules are usually subject to disruptions which
have to be repaired by reactive scheduling decisions. The most popular
predictive approach in project management and machine scheduling
literature is to leave idle times (time buffers) in schedules in coping
with disruptions, i.e. the resources will be under-utilized. Therefore,
preparing initial schedules by considering possible disruption times
along with rescheduling objectives is critical for the performance of
rescheduling decisions. In this paper, we show that if the processing
times are controllable then an anticipative approach can be used to
form an initial schedule so that the limited capacity of the production
resources are utilized more effectively. To illustrate the anticipative
scheduling idea, we consider a non-identical parallel machining
environment, where processing times can be controlled at a certain
compression cost. When there is a disruption during the execution of
the initial schedule, a match-up time strategy is utilized such that a
repaired schedule has to catch-up initial schedule at some point in
future. This requires changing machine-job assignments and processing
times for the rest of the schedule which implies increased
manufacturing costs. We show that making anticipative job sequencing
decisions, based on failure and repair time distributions and
flexibility of jobs, one can repair schedules by incurring less
manufacturing cost. Our computational results show that the match-up
time strategy is very sensitive to initial schedule and the proposed
anticipative scheduling algorithm can be very helpful to reduce
rescheduling costs. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1711.txt#


Over the past few years, we have been developing techniques for
high-speed 3D shape measurement using digital fringe projection and
phase-shifting techniques: various algorithms have been developed to
improve the phase computation speed, parallel programming has been
employed to further increase the processing speed, and advanced
hardware technologies have been adopted to boost the speed of
coordinate calculations and 3D geometry rendering. We have successfully
achieved simultaneous 3D absolute shape acquisition, reconstruction,
and display at a speed of 30 frames/s with 300 K points per frame. This
paper presents the principles of the real-time 3D shape measurement
techniques that we developed, summarizes the most recent progresses
that have been made in this field, and discusses the challenges for
advancing this technology further. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1757.txt#


Consider the problem of scheduling a set of jobs to be processed
exactly once, on any machine of a set of unrelated parallel machines,
without preemption. Each job has a due date, weight, and, for each
machine, an associated processing time and sequence-dependent setup
time. The objective function considered is to minimize the total
weighted tardiness of the jobs.This work proposes a non-delayed
relax-and-cut algorithm, based on a Lagrangean relaxation of a time
indexed formulation of the problem. A Lagrangean heuristic is also
developed to obtain approximate solutions.Using the proposed methods,
it is possible to obtain optimal solutions within reasonable time for
some instances with up to 180 jobs and six machines. For the solutions
for which it is not possible to prove optimality, interesting gaps are
obtained. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1771.txt#


In this paper we consider a network of distributed sensors which
simultaneously measure a physical parameter of interest, subject to a
certain probability of sensing error. The sensed information at each of
such nodes is channel-encoded and forwarded to a central receiver
through parallel independent AWGN channels. In this scenario, several
recent contributions have shown that the end-to-end Bit Error Rate
(BER) performance can be dramatically improved if the decoders
associated to each received signal and the data fusion stage exchange
soft information in an iterative Turbo-like fashion. In order to
achieve optimum performance, the probability of sensing error must be
known (or estimated) at the receiver. In this work we describe a novel
method for estimating such sensing error probability by properly
weighting likelihoods output from the Soft-Input Soft-Output decoders
(SISO), which is shown to outperform other estimation methods based in
hard-decision comparisons, specially in the low

#e1#

#s1#
#abstract1772.txt#


The following topics are dealt with: discrete chaotic system; variable
structure control; networked control system; image quality assessment;
ant colony algorithm; wireless sensor networks; remote sensing image
segmentation; BP neural network; blind source seperation; face
detection; skin color segmentation; flight vehicle; secret sharing;
fault tree analysis; humanoid robot; adaptive feedforward control;
mobile robot; video transmission system; vibration control; trajectory
planning method; image processing; unmanned surface vehicles;
differential geometry method; fault diagnosis; embedded system; spatial
parallel manipulator; fuzzy control algorithm; wavelet analysis
denoising; evolutionary programming; genetic algorithm; time delayed
control systems; particle swarm optimisation; flight management;
underwater vehicles; torque control; servo drive system; hydraulic
power steering system; PID control; digital watermarking; handwriting
signature verification; semantic programming language; CNC system;
agent-based modeling.

#e1#

#s1#
#abstract1773.txt#


In this paper, a new method combined fuzzy theory and neural network
was proposed, which was employed to solve multi-maneuvering target
tracking. Multi-maneuvering target tracking is a process, which manages
measure information to maintain currently state estimate. The fields of
fuzzy sets and neural network have made rapid progress in recent years.
Neural networks are nonlinear network of self-organization and
self-learning, they possess the capabilities of large-scale parallel
processing, distributed information, neural network have important
influence on the revolution of traditional target tracking theory. The
training of fuzzy rule-based systems by using the learning ability of
neural network can improve the expression ability of the network.
Combining the fuzzy theory and neural network is a new approach to
solve multi-maneuvering target tracking.

#e1#

#s1#
#abstract1774.txt#


In parallel database system, a good data placement could improve
execution efficiency of multi-join queries greatly. The bandwidth of
network communication is always the bottleneck of parallel database
system based on PC clusters. Data communication among nodes would bring
more time cost when executing join operations. This paper proposes
selection of nodes algorithm, which takes the data redistribution into
consideration and reduces additional communication cost. Furthermore,
it takes into account intra-operator parallelism, independent
inter-operator parallelism and pipelined parallelism in order to
develop parallelisms of PC clusters system. The result of experiment
indicates the algorithm has good performance and contributes to
promoting execution efficiency of parallel multi-join queries.

#e1#

#s1#
#abstract1775.txt#


Partially parallel imaging with localized sensitivities is a fast
parallel image reconstruction method for both Cartesian and
non-Cartesian trajectories, but suffers from aliasing artifacts when
there are deviations from the assumption of perfect localization. Such
reconstructions would normally crop the individual coil images to
remove the artifacts prior to combination. However, the sampling
densities in variable-density /i k/-space trajectories support
different field-of-views for separate regions in /i k/ -space. In fact,
the higher sampling density of low frequencies can be used to
reconstruct a bigger field-of-view without introducing aliasing
artifacts and the resulting image signal-to-noise ratio (

#e1#

#s1#
#abstract1776.txt#


Multi-Input Multi-Output (MIMO) wireless communication systems commonly
employ beamforming techniques with Singular Value Decomposition (SVD).
In such systems, if no channel encoding is employed, the full diversity
order provided by the channel is achieved when a single symbol is
transmitted over multiple channels; however, this property is lost
whenever multiple symbols are simultaneously transmitted. The full
diversity order can be restored when channel coding is added to such a
system. For example, when Bit-Interleaved Coded Modulation (BICM) is
combined with this technique, the full diversity order of NM in an M *
N MIMO channel, transmitting S parallel streams is possible; provided
SR/sub c/ #or= 1 where R/sub C/ is the BICM convolutional code rate. In
this paper, we present multiple beamforming with constellation
precoding which can achieve the full diversity order with both uncoded
and BICM-coded SVD systems. An analytical proof of this property is
provided. In addition, to reduce the computational complexity of
Maximum Likelihood (ML) decoding, we introduce a Sphere Decoding (SD)
technique. This technique achieves several orders of magnitude
reduction in computational complexity not only with respect to
conventional ML decoding, but also, with respect to conventional SD.

#e1#

#s1#
#abstract1777.txt#


The development of automatic guided wave interpretation for detecting
corrosion in aluminum aircraft structural stringers is described. The
dynamic wavelet fingerprint technique (DWFP) is used to render the
guided wave mode information in two-dimensional binary images.
Automatic algorithms then extract DWFP features that correspond to the
distorted arrival times of the guided wave modes of interest, which
give insight into changes of the structure in the propagation path. To
better understand how the guided wave modes propagate through real
structures, parallel-processing elastic wave simulations using the
elastodynamic finite integration technique (EFIT) has been performed.
3D simulations are used to examine models too complex for analytical
solutions. They produce informative visualizations of the guided wave
modes in the structures, and mimic the output from sensors placed in
the simulation space. Using the previously developed mode extraction
algorithms, the 3D EFIT results are compared directly to their
experimental counterparts.

#e1#

#s1#
#abstract1778.txt#


Field modeling is a common practice in NDE. However, it is a very time
consuming task, making difficult its use in systems demanding real time
response. The use of Graphics Processing Units (GPUs), which are
programmable devices with a high level of parallelism and a good
relation between price and performance can accelerate the computations.
This work shows that using GPU technology the computing time is reduced
in more than one order of magnitude with respect to CPU implementations.

#e1#

#s1#
#abstract1779.txt#


This paper presents an efficient stiffness identification technique for
truss structures based on distributed local computation. Sensor nodes
on each element are assumed to collect strain data and communicate only
with sensors on neighboring elements. This can significantly reduce the
energy demand for data transmission and the complexity of transmission
protocols, thus enabling a simplified wireless implementation. Element
stiffness parameters are identified by simple low order matrix
inversion at a local level, which reduces the computational energy,
allows for distributed computation and makes parallel data processing
possible. The proposed method also permits addressing the problem of
missing data or faulty sensors. Numerical examples, with and without
missing data, are presented and the element stiffness parameters are
accurately identified. The computation efficiency of the proposed
method is n/sup 2/ times higher than previously proposed global damage
identification methods.

#e1#

#s1#
#abstract1780.txt#


An integrated simulation method for investigating nonlinear sound beams
and 3D acoustic scattering from any combination of complicated objects
is presented. A standard finite-difference simulation method is used to
model pulsed nonlinear sound propagation from a source to a scattering
target via the KZK equation. Then, a parallel 3D acoustic simulation
method based on the finite integration technique is used to model the
acoustic wave interaction with the target. Any combination of objects
and material layers can be placed into the 3D simulation space to study
the resulting interaction. Several example simulations are presented to
demonstrate the simulation method and 3D visualization techniques. The
combined simulation method is validated by comparing experimental and
simulation data and a demonstration of how this combined simulation
method assisted in the development of a nonlinear acoustic concealed
weapons detector is also presented.

#e1#

#s1#
#abstract1781.txt#


We describe the use of ultrasonic guided waves for identifying the mass
loading due to underwater limpet mines on ship hulls. The Dynamic
Wavelet Fingerprint (DFWP) technique is used to render the guided wave
mode information in two-dimensional binary images because the waveform
features of interest are too subtle to identify in time domain. The use
of wavelets allows both time and scale features from the original
signals to be retained, and image processing can be used to
automatically extract features that correspond to the arrival times of
the guided wave modes. For further understanding of how the guided wave
modes propagate through the real structures, a parallel processing, 3D
elastic wave simulation is developed using the finite integration
technique. This full field technique models situations that are too
complex for analytical solutions, such as built-up 3D structures. The
simulations have produced informative visualizations of the guided wave
modes in the structures as well as mimicking directly the output from
sensors placed in the simulation space for direct comparison to
experiments. Results from both drydock and in-water experiments with
dummy mines are also shown.

#e1#

#s1#
#abstract1782.txt#


A new two-dimensional (2D) holographic microwave imaging technique is
proposed to reconstruct the 2D image of a target. It is based on the
Fourier analysis of the data recorded by two antennas scanning together
two separate rectangular parallel apertures on both sides of a target.
The complex back-scattered signals of the two antennas are first
processed to localize the target in the range direction. Then, the 2D
image of the target is reconstructed. No assumptions are made about the
incident fields, which can be derived by either simulation or
measurement. Both the back-scattered and forward-scattered signals can
be used to reconstruct the image of the target. This makes the proposed
technique applicable with near-field measurements. To evaluate the
proposed technique, the range localization and the 2D image
reconstruction of a predetermined simulated target are examined.
Associated resolution limits, sampling constraints and the impact of
noise are also discussed.

#e1#

#s1#
#abstract1783.txt#


As a practical scheme carrying out the high performance computing tasks
for small research groups and personal researchers, we propose a small
Beowulf PC-cluster parallel system based on MPI and Windows by
employing a set of PCs and 100 M Ethernet. The experimental evaluations
on the classical case of computing pi show that the system has
satisfactory speedup and parallel efficiency. Owing to the 100 M
Ethernet which has less bandwidth and more communication delay, the
speedup and parallel efficiency will come down with increasing working
nodes, and a more excellent high speed network should be build.

#e1#

#s1#
#abstract1784.txt#


In order to share the intelligence among all kinds of mobile nodes,
this paper puts forward the intelligence sharing technology based on
mobile agents in mobile peer-to-peer (MP2P) system. The MP2P system
could build up a common applicable theoretical framework for mobile
intelligence management such as mobile intelligence acquisition, query,
synchronization, cache, and prefetching. In order to reduce the
maintenance cost of intelligence index table for the mobile nodes in
the MP2P network, use the edge intelligence of network enough, and
parallel process the mobile task of network, we propose the
intelligence management technology based on mobile agent (MA), which
improves the execution efficiency of security authentication,
intelligence distribution, cache prefetching, transaction processing,
location service, and task cooperation among the mobile nodes.

#e1#

#s1#
#abstract1785.txt#


By taking the process of synthetic ammonia decarbornization as the
research object, a new method of early real-time fault diagnosis based
on the linear classifier-reforming neural network was proposed. The
method, which need not establish accurate mathematical model, and has
the advantages of its simple learning algorithm, accumulate knowledge
from example automatically, learning and classification of parallel
processing and fast response speed etc. The results show that it can be
applied to early real-time fault diagnosis in the process, and can
provide techniques guarantee for safety production.

#e1#

#s1#
#abstract1786.txt#


As the demand for global information increases significantly,
multilingual corpora has become a valuable linguistic resource for
applications to cross-lingual information retrieval and natural
language processing. A Web-based English-Chinese bilingual parallel
corpus of automatic Construction Technology solved the shortage of
bilingual English-Chinese Parallel Corpus. First, some web pages which
may be set translation dig of from a particular source, and then from
the web pages focused on the external characteristics according to the
similarity to extract the candidate web pages in parallel pairs, use of
content-based methods on parallel web pages for each of these
candidates assessed. In the assessment of the candidate pairs of
parallel web pages, this paper design ECVS models of bilingual text
similarity assessed based on the classic vector space model.

#e1#

#s1#
#abstract1787.txt#


Reducing dimension processing is needed in feature samples because the
repeated and secondary features would reduce the classification ability
and increase computation complexity. In this paper, a feature selection
method, named MPSO (Modified Particle Swarm Optimization), is proposed.
The original group velocity of a particle swarm was changed into two
separate and parallel particle swarm velocity, which was effectively
and quickly applied to the feature extraction of the optimum samples on
the basis of Discrete Binary PSO. Then the least squares support vector
machine classifier is used to verify the feasibility of this method.
The experimental results show that, compared with the method in the
literature, the iteration times in this method are only 17 times in
average, while the iteration times in the literature are 23 times; the
selected features and the average recognition accuracy after feature
selection are slightly better than the ones in the method in the
literature. Therefore,the proposed method is feasible and effective.

#e1#

#s1#
#abstract1788.txt#


A new analog-to-digital converter (ADC) technology called Single-slope
look ahead ramp (SSLAR) analog-to-digital converter (ADC) was proposed
for column-parallel CMOS image sensors. Additionally, a corresponding
programmable ramp generator for SSLAR ADC was also designed in such a
way that it only allows flexible code hopping (between 0 and 127 least
significant bit (LSB)), code fall back and look-ahead operations in
column-parallel ramp ADC. This new ADC technology is able to provide
conversion speed improvement depending on individual image information
with less than 1.0% image quality degradation. Simulation demonstrated
that the conversion speed of this new ADC technology is 4-5x higher
than a traditional single-slope ADC with minimal circuitry for
processing 8-bit standard gray images as well as 3-4x for standard CIF
videos. For processing higher resolution images, the conversion speed
may further increase. A prototype chip using the 8-bit SSLAR ADC
architecture was realized in a 0.5 mu m, 2P3M, CMOS process with a
layout area of 8.2 mm/sup 2/.

#e1#

#s1#
#abstract1792.txt#


Bilingual parallel corpora, also known as bitexts, convey the same
information in two different languages. This implies that to model a
bitext we can take advantage of the translation relationship that
exists between the two texts; the text alignment task makes it possible
to establish such a translation relationship. A biword is defined as a
pair of words, each from a different text, that are mutual translations
in the bitext; the use of biwords allows both texts in the bitext to be
represented on a single model. Several biword-based schemes have been
proposed leading to good compression ratios. Bearing in mind Melamed's
affirmation which states that "the translation of a text into another
language can be viewed as a detailed annotation of what that text
means", we propose a new model for bitexts in agreement with this
affirmation, dubbed MAR. The idea is to represent the words in the
right text with respect to the preceding word in the left text; thus, a
first-order model based on alignment relationships is proposed.

#e1#

#s1#
#abstract1793.txt#


This work proposes a novel practical and general-purpose lossless
compression algorithm named Neural Markovian Predictive Compression
(NMPC), based on a novel combination of Bayesian Neural Networks (

#e1#

#s1#
#abstract1794.txt#


Electron tomography (ET) allows elucidation of the three-dimensional
(3D) structure of large complex biological specimens at molecular
resolution. In order to achieve such resolution levels, large
projection images have to be used to compute the 3D reconstructions.
Tomographic reconstruction on this scale requires a tremendous use of
computational resources and a considerable processing time. In this
work, we present and evaluate a highly optimized implementation of the
Weighted Back-Projection reconstruction algorithm. Briefly,
optimizations made to the code comprise (1) vector processing with SSE
(Streaming SIMD Extensions) instructions, (2) an efficient use of cache
memory, (3) to take advantage of the inherent image symmetry, (4) to
use the FFTW (Fastest Fourier Transform in the West) library for image
filtering, (5) to use regions of interest and last, but not least, (6)
a wide range of minor optimizations like some data pre-calculations or
an instruction level parallelism improvement. We have evaluated the
method on tomographic reconstructions of several datasets and on two
computing platforms. The results show that our version speeds up the
method by a factor around 14 or 16, depending on the platform.

#e1#

#s1#
#abstract1795.txt#


High performance computing is becoming critical in the medical area to
aid real-time processing of complex analysis of biological signals. In
this paper parallel schemes for real-time computations of pair-wise
correlation (PWC) of electroencephalogram (EEG) signals, which belongs
to streaming-data class of applications, are proposed and implemented
and their performances are evaluated. Currently most of the EEG based
diagnosis for epilepsy is done off-line. However, there is a growing
need to perform these diagnoses in real-time to aid health care
providers, including surgeons, in decision-making process that will
lead to improved quality of life and prevent undesirable consequences,
such as readmission to hospitals resulting in prolonged suffering and
higher health care costs. Systematic study of the PWC problem and the
IBM Cell Broadband Engine (CBE) architecture led us to a model that is
well suited for the Cell architecture and GPUs. Measurements on the CBE
indicate that speedup of 33.91 is possible over the serial code running
on Intel Xeon processor and the schemes can be used for real-time
signal processing.

#e1#

#s1#
#abstract1796.txt#


A new gamma-camera architecture named HiSens is presented and
evaluated. It consists of a parallel hole collimator, a pixelated
CdZnTe (CZT) detector associated with specific electronics for 3D
localization and dedicated reconstruction algorithms. To gain in
efficiency, a high aperture collimator is used. The spatial resolution
is preserved thanks to accurate 3D localization of the interactions
inside the detector based on a fine sampling of the CZT detector and on
the depth of interaction information. The performance of this
architecture is characterized using Monte Carlo simulations in both
planar and tomographic modes. Detective quantum efficiency (DQE)
computations are then used to optimize the collimator aperture. In
planar mode, the simulations show that the fine CZT detector
pixelization increases the system sensitivity by 2 compared to a
standard Anger camera without loss in spatial resolution. These results
are then validated against experimental data. In SPECT, Monte Carlo
simulations confirm the merits of the HiSens architecture observed in
planar imaging.

#e1#

#s1#
#abstract1797.txt#


A scheme of pulse based neural circuits for object tracking is
proposed. Different from the conventional frame-based methodology, the
proposed design utilises parallel arrays of circuits to extract pixels
with significant temporal contrast, which are indicators of moving
objects. It can be implemented on an FPGA chip in full parallelism and
improves the tracking performance dramatically. Moreover, its
integrate-and-fire neural model can cluster concave data sets.
Experimental results show that the proposed scheme outperforms
conventional methods.

#e1#

#s1#
#abstract1798.txt#


Load balancing algorithms are an essential component of parallel
computing reducing the response time of applications. Frequently,
balancing algorithms have a centralized behavior requiring a lot of
messages to operate, thus causing scalability problems. A solution to
improve scalability is to define a decentralized algorithm, avoiding
the generation of bottlenecks. DLML (Data List Management Library) is a
tool that, in a transparent way, allows the parallel processing of data
that are organized through a List. One drawback of this tool is the
global bidding algorithm used to distribute the data (work) generated
during the execution. In this paper two load balancing algorithms for
DLML handling partial information are proposed. The first algorithm
considers a logical Torus topology and the second one follows a Binary
Tree topology for communications. Results show how the scalability of
DLML was improved, using two clusters of 40 and 1024 processing units,
and executing dynamic and static applications.

#e1#

#s1#
#abstract1799.txt#


A new method for obtaining models of the performance of parallel
applications based on statistical analysis is presented in this paper.
This method is based on the Akaike's information criterion (AIC) that
provides an objective mechanism to rank different models by means of an
experimental data fit. The input of the modeling process is a set of
variables and parameters that can a priori influence the performance of
the application. This set can be provided by the user. Using this
information, the method automatically generates a set of candidate
models. These models are fit to the experimental data and the AIC score
of each model is calculated. The model with the best AIC score is
selected as the best model. Also, using the AIC scores of all candidate
models, useful statistical information is provided to help the user to
evaluate the quality of the selected model, as well as indications of
how to interactively improve this modeling process. As a first case of
study, statistical models obtained for different implementations of the
broadcast collective communication in Open MPI are shown. These models
are very accurate, exceeding its adjustment to theoretical approaches
based on the LogGP model. Finally, the NAS Parallel Benchmark is also
characterized using this new method with good results in terms of
accuracy.

#e1#

#s1#
#abstract1800.txt#


Network security applications such as to detect malware, security
breaches, and covert channels require packet inspection and processing.
Performing these functions at very high network line rates and low
power is critical to safe guarding enterprise networks from various
cyber-security threats. Solutions based on FPGA and single or
multi-core CPUs has several limitations with regards to power and the
ability to match the ever increasing line rates. This paper describes a
MPPA (Massively Parallelized Processing Architecture) framework based
on the Ambric parallel processing device that can speed up computation
of network packet processing and analysis tasks. This is accomplished
with a programmable processor interconnection that enables
parallelizing the application and replication of data through channels.
In this paper, we consider three network security applications -
detecting malware, detecting covert timing channels, and a symmetric
encryption engine. Experimental analyses of parallel implementations of
the detection algorithms show that MPAA can easily achieve throughput
greater than 1 Gbps with low power usage.

#e1#

#s1#
#abstract1801.txt#


Defining performance models associated with the application structure
has been proven a useful strategy for implementing dynamic tuning
tools. However, for extending this strategy to more complex
applications (those composed by different structures) it must integrate
a policy for the distribution of the resources among the different
application components. Consequently, we propose to take advantage of
the knowledge of these models and combine them with a resource
management policy for obtaining a global model. In this sense, this
work constitutes the ongoing effort in the development of performance
models for dynamic tuning.

#e1#

#s1#
#abstract1802.txt#


This paper presents the first experimental results of the use of our
new adaptive tool for synchronization, based on ordered read-write
locks, ORWL. They provide a new synchronizing method for data-oriented
parallel algorithms and are particularly suited for iterative pipelined
algorithms with out-of-core data. We conducted experiments with the
classic benchmarking Livermore Kernel 23 algorithm to validate the
theoretical model and measure the efficiency of the first available
implementation of ORWL in the PARXXL library. They show that this tool
is able to efficiently control an IO bound application running on 64
parallel POSIX threads with tight data dependencies between them.

#e1#

#s1#
#abstract1803.txt#


The shift to multicore processors demands efficient parallel
programming on a diversity of architectures, including homogeneous and
heterogeneous chip multiprocessors (CMPs). Task parallel programming is
one approach that maps well to CMPs. In this model, the programmer
focuses on identifying parallel tasks within an application, while a
runtime system takes care of managing, scheduling, and balancing the
tasks among a number of processors or cores. Heterogeneous CMPs, such
as the Cell Broadband Engine, present new challenges to task parallel
programming and corresponding runtime systems. In this paper, we
present a library based on task pools for dynamic task scheduling and
load balancing on Cell processors. In contrast to other approaches, our
task pools include support for creating tasks using the Synergistic
Processing Elements (SPEs), which enables the implementation of a wide
range of task parallel applications. Our experiments show that task
pools provide flexible and efficient support for task parallel
programming on Cell processors. In addition, we show that offloading
the process of task creation from the PPE to the SPEs provides much
potential for exploiting fine-grained parallelism.

#e1#

#s1#
#abstract1804.txt#


The following topics are dealt with: scheduling; resource management;
fault tolerance; performance modeling; cloud computing; networking;
multi-many core systems; programming abstractions; scalability;
peer-to-peer environments; next generation Web computing;
bioinformatics; grid computing; high performance computing; nuclear
fusion applications; on-chip parallel system; network-based systems;
distributed systems; parallel algorithms; sparse linear algebra
computations; network security and distributed systems security.

#e1#

#s1#
#abstract1805.txt#


We present in this paper a novel method to predict application runtimes
on backfilling parallel systems. The method is based on mining
historical data to obtain important parameters. These parameters are
then applied to predict the runtime of future applications. It has been
shown in previous works that both underestimate and inaccuracy in
prediction have adverse impacts on scheduling performance of
backfilling systems. In our study, we try to reduce the number of jobs
that are underestimated and reduce the prediction error as much as
possible. Comparing with other predictors, experimental results show
that our predictor is up to 25% better with respect to the problem of
underestimate. Moreover, using the metric proposed in for the accuracy,
our predictor improves up to 32%.

#e1#

#s1#
#abstract1806.txt#


Pervasive Grid Computing Platforms include centralized computing nodes
(e. g. parallel servers) as well as decentralized and mobile devices.
Pervasive Grid applications include data- and computing-intensive
components which can be mapped also onto decentralized and mobile
nodes. The effective and practical success of this mapping resides also
in deriving proper configurations of applications which consider the
limited memory capabilities of those resources. In this paper we target
this issue by showing how we can study and configure the memory
requirements of an Emergency Management application. We present our
solutions by using the ASSISTANT programming model for Pervasive Grid
applications.

#e1#

#s1#
#abstract1807.txt#


Putting performance asymmetric cores inside the same processor can be a
good alternative to obtain high performance per area, throughput and
single-threaded performance. However, the impact of running parallel
applications on this type of machine is not clear, since most of
previous work focused on multi-programmed and server workloads where
there is low or no dependence between threads. In this work, we analyze
the impact of running parallel shared-memory programs on heterogeneous
multi-core setups using six parallel applications with diverse
parallelization schemes. Moreover, we show that, in some cases, with a
high number of cores, it is better to put one complex core than several
simple ones. The impact of sharing the address space between asymmetric
cores with private caches was also investigated and the number of
invalidations per write access was not greater than a comparable
homogeneous configuration.

#e1#

#s1#
#abstract1808.txt#


An emerging application field for structure matching is related to in
silico studies of molecular biology: considering that protein function
is mainly related to its external morphology, the possibility to match
macromolecular surfaces is very important to infer information about
the interaction of biological components. In this work we present a
parallel algorithm based on images of local description, originally
designed for object recognition in robotics, which has been adapted to
identify surface complementarities for screening the interaction
possibilities of biological macromolecules. The results obtained show
the good discrimination power of the proposed approach, coupled with
high performance figures.

#e1#

#s1#
#abstract1809.txt#


Tissue MicroArray technology aims to perforin inimunohistocheniical
staining on hundreds of different tissue samples simultaneously,
allowing faster analysis and considerably reducing costs incurred in
staining. The presented work supports the pre-array phase of this
technique, i.e. the automatic discrimination between normal and
pathological regions within the analyzed tissues, and it works in the
specific context of tubular breast cancer. The diagnosis is performed
by automatically analyzing specific morphological features of the
breast samples, in order to define if tissues present a normal behavior
either they show pathological characteristics, in particular the
absence of a double layer of cells around the lumen or the decay of a
regular glands-and-lobules structure. Tissue structure is investigated
through a parallel image processing algorithm, which performs the
extraction of morphological parameters from the acquired images and
compares them to experimentally validated threshold values. The input
image is divided into a number of independent sub-images, to be
singularly analyzed and labeled according to the result of the computed
diagnosis. The time spent to actually analyze each sub-image generally
varies depending on the pathological characteristics of each part of
the tissue. In order to properly manage and exploit this feature of the
algorithm, the analysis of the sub-images is dynamically dispatched
among parallel processes. Experimental results, carried on by
exploiting the parallel paradigm, certify the actual improvement in the
execution time, leading to almost linear speed-up values on the actual
tissue elaboration.

#e1#

#s1#
#abstract1810.txt#


Current wide availability of multicore systems requires tools that can
help scientists to smoothly update their applications to take advantage
of the parallel processing capabilities of these systems. In this
paper, we present an experience with aspect-oriented programming (AOP)
techniques to perform this move. We describe the parallelization of a
Java library that implements algorithms from the Evolutionary
Computation field (JECoLi), applied to two case studies in
Bioinformatics, namely the optimization of feeding profiles in
fed-batch fermentations and in silico strain optimization in Metabolic
Engineering. AOP allowed us to enable the library to take advantage of
multicore systems with minimal impact on the original code and to
simultaneously develop the parallelization and the original library.
Moreover, we developed modules that extend the library's behavior for a
better usage of multicore resources. Performance results show that this
approach boosts performance, does not compromise the quality of the
final solutions and enables a more loosely coupled development.

#e1#

#s1#
#abstract1811.txt#


Coupling separately developed codes offers an attractive method for
increasing the accuracy and fidelity of the computational models.
Examples include the earth sciences and fusion integrated modeling.
This paper describes the Framework Application for Core-Edge Transport
Simulations (FACETS).

#e1#

#s1#
#abstract1812.txt#


One of the central tasks of EUFORIA is to port, parallelise, and
optimise fusion simulation codes, developed at individual research
institutes in Europe. There are three supercomputer centres involved in
the project located at Barcelona, Edinburgh, and Helsinki. For some of
the fusion codes simply porting them to one of the supercomputers
represents a major advancement in the use of the codes, as they until
now have mainly been used by a small user community, or even
exclusively by the author of the code. Also, where codes currently can
only use one processor (i. e. are serial) providing any parallel
functionality can be of major benefit to the code and the code
owner(s). Many of the simulation codes for edge and core transport
modelling of fusion plasma using high performance computing are
estimated to currently require weeks or months of execution time to
simulate science at a scale required to model the new fusion reactor
ITER, and therefore these codes have to be optimised to run as fast as
possible and parallelised in such a way that computer resources are
used as effectively as possible. During the first fifteen month of the
project, we have successfully ported eleven fusion codes to the
supercomputers in Barcelona, Edinburgh and Helsinki. The installation
procedure, library requirements and runtime scripts have been
documented for each code, and deposited in the EUFORIA software
repository and code revision system. Following this a number of these
codes have been chosen for code optimisation and improvements in
parallelisation and this paper outlines the experience that we have had
with some of these codes, the performance improvements achieved, and
the techniques used.

#e1#

#s1#
#abstract1813.txt#


In this study, we developed a high speed eigenvalue solver that is the
necessity of plasma stability analysis system for International
Thermo-nuclear Experimental Reactor (ITER) on Cell cluster system. Our
stability analysis system is developed in order to prevent damages to
the ITER from plasma disruption. MARG2D, which is the main part of the
system, analyzes the state of plasma with the measured conditions
instantaneously. According to our estimation, the most time consuming
part of MARG2D is eigensolver and we must solve resulted eigensystem
whose dimension is hundred thousand within a second. However, current
massively parallel processor (MPP) type supercomputer is not applicable
for such instantaneous calculation, because the overhead of network
communication becomes dominant. Therefore, we employ Cell cluster
system, whose processor has higher performance than MPP's, because we
can obtain sufficient processing power with small number of processors.
Furthermore, we developed novel eigenvalue solver with the
consideration of hierarchical architecture of Cell cluster: Inter
processor, intra processor, and SIMD parallelism. Finally, we succeeded
to solve the block tridiagonal Hermitian matrix, which had 1024
diagonal blocks and the size of each block was 128 * 128 within a
second.

#e1#

#s1#
#abstract1814.txt#


Media spaces provide users with flexible support for easy interaction
with technology and with each other, both at the same place and over
distance. From a technological perspective the development of these
environments is often inefficient, since most environments are
developed specifically, without any synergies or reuse of previous
concepts and implementations. In this paper we present the cooperative
media space PPPSpace that is based on powerful technical parallel and
distributed software engineering concepts and at the same time easy to
use for end-users.

#e1#

#s1#
#abstract1815.txt#


We present a parallel conjugate gradient solver for the Poisson problem
optimized for multi-GPU platforms. Our approach includes a novel
heuristic Poisson preconditioner well suited for massively-parallel
SIMD processing. Furthermore, we address the problem of limited
transfer rates over typical data channels such as the PCI-express bus
relative to the bandwidth requirements of powerful GPUs. Specifically,
naive communication schemes can severely reduce the achievable speedup
in such communication-intense algorithms. For this reason, we employ
overlapping memory transfers to establish a high level of concurrency
and to improve scalability. We have implemented our model on a
high-performance workstation with multiple hardware accelerators. We
discuss the mathematical principles, give implementation details, and
present the performance and the scalability of the system.

#e1#

#s1#
#abstract1816.txt#


In many numerical applications resulting from computational science and
engineering problems, the solution of sparse linear systems is the most
prohibitively compute intensive task. Consequently, the linear solvers
need to be carefully chosen and efficiently implemented in order to
harness the available computing resources. Krylov subspace based
iterative solvers have been widely used for solving large systems of
linear equations. In this paper, we focus on the design of such
iterative solvers to take advantage of massive parallelism of general
purpose Graphics Processing Units (GPU)s. We will consider Stabilized
BiConjugate Gradient (BiCGStab) and Conjugate Gradient Squared (CGS)
methods for the solutions of sparse linear systems with unsymmetric
coefficient matrices. We discuss data structures and efficient
implementation of these solvers on the NVIDIA's CUDA platform. We
evaluate scalability and performance of our implementations in the
context of a financial engineering problem of solving multidimensional
option pricing PDEs using sparse grid combination technique.

#e1#

#s1#
#abstract1817.txt#


With fast development of transistor technology, Graphic Processing
Unit(GPU) is increasingly used in the non-graphics applications, and
major GPU hardware vendors have introduced software stacks for their
own GPUs, such as Brook+ for AMD GPU. Compared with the traditional
parallel systems, heterogeneous systems integerating stream-based
multi-threaded GPUs provide higher parallel computing capabilities with
lower cost. However, porting traditional applications to the
heterogeneous systems makes new demand of application optimization on
GPU. Based on the AMD's Brook+ platform, we explored application
optimization features on AMD GPU by optimizing and implementing the
benchmark LBM from SPEC2006. To improve the program locality, we
optimized the original data layout of LBM. Using the short vector data
types mechanism provided by Brook+, we also optimized the GPU's
bandwidth utilization and its thread processors' efficiency. Through
the branch elimination technique, we reduced the performance lose
caused by branch divergences in the kernel, which is due to the GPU's
SIMD executing mode. The experiment results show that data layout,
memory bandwidth, branch paths and other factors have a close effect on
the performance of program execution on the GPU. Through all the
optimizations, we finally got a speedup of 22x (single-precision) and
19x (double-precision) over the original serial benchmark code on a
Quad-core CPU, and a speedup of 4x (single-precision) and 8.7x
(double-precision) over the original OMP benchmark code on a 8-core CPU.

#e1#

#s1#
#abstract1818.txt#


In order to effectively handle the growing amount of available RDF
data, a scalable and flexible RDF data processing framework is needed.
We previously proposed a Hadoop-based framework, which takes advantages
of scalable and fault-tolerant distributed processing technologies,
originally proposed as Google's distributed file system and MapReduce
parallel model. In this paper, we present a method extending the Pig
data processing platform on top of the Hadoop infrastructure. Pig
compiles programs written in a high level language, called Pig Latin,
into MapReduce programs that can be executed by Hadoop. In order to
support RDF, Pig was extended with the ability to load and store RDF
data efficiently. Furthermore, as reasoning is an important requirement
for most systems storing RDF data, support for inferring new triples
using entailment rules was also added. In this paper, we describe these
extensions and present an evaluation of their performance.

#e1#

#s1#
#abstract1819.txt#


In the context of service composition and orchestration, service
invocation is typically scheduled according to execution plans, whose
topology establishes whether different services are to be invoked in
parallel or in a sequence. In the latter case, we may have a
configuration, called pipe join, in which the output of a service is
used as input for another service. When the services involved in a pipe
join output results sorted by score, the problem arises of efficiently
determining the join tuples (aka combinations) with the highest
combined scores. In this paper we study different execution strategies
related to the pipe join configuration. First, we consider a strategy
that minimizes the access costs to achieve a target number of
combinations. Then, we propose a strategy that explicitly considers the
scores of the output tuples in order to provide deterministic
guarantees that the top-k combinations have been found. Finally, a
hybrid strategy is presented.

#e1#

#s1#
#abstract1820.txt#


Many organizations are aiming to move away from traditional batch
processing ETL to real-time ETL (RT-ETL). This move is motivated by a
need to analyze and take decisions on as fresh a data as possible. The
RT-ETL engines operate on the abstraction of data flow executed on
parallel architectures. For high throughput and low response times,
there is a need for partitioning the data over the large number of
nodes in the engine. In this paper, we consider the problem of
partitioning realtime ETL flows and we propose a high level
architecture for that.

#e1#

#s1#
#abstract1821.txt#


Designing efficient cache, memory, and storage subsystem for modern
embedded systems supporting a variety of applications is a great need.
Embedded systems are being deployed with multicore processors to help
parallel and distributed computing in order to meet the requirements
for increased processing speed. Multiple cores offer manifold options
to organize multi-level caches. A mixture of cache memory hierarchies
are proposed to satisfy the requirements of high-performance low-power
multicore embedded systems. In this paper, we investigate the impact of
CL2 organizations on the performance and power consumption for
multicore embedded systems. We simulate two 4-core architectures, one
with shared CL2 and the other one with private CL2s. We use MPEG4, FFT,
MI, and DFT applications/algorithms in our experiment. Simulation
results depict that the mean delay and total power consumption
significantly vary with the variations of CL2 organization and
applications. It is observed that reductions in total power consumption
and mean delay per task of up to 43% and 36%, respectively, are
possible with optimized CL2, with an optimal choice of 256 KB CL2
cache, 64 B CL2 line size, and 8-way CL2 associativity level.

#e1#

#s1#
#abstract1822.txt#


A neuronal recording system for brain-machine interfaces (BMI) based on
asynchronous biphasic pulse coding is described. A recording experiment
comparing, in parallel, a commercial recording system (Tucker-Davis
Technology) and the UF's custom solution (FWIRE) is set up to compare
performance. The novel aspect of the UF system is that the analog
signal is represented by an asynchronous pulse train, which provides a
low-power, low-bandwidth, noise-resistant means for coding and
transmission. Based on different front-end hardware settings, recording
bandwidth and corresponding reconstruction accuracy can be varied.
Taking advantage of neural firing features, the pulse-based approach
requires less than 3 K pulses/second to record a 25 KHz bandwidth
signal from a hardware neural simulator. Recording performance has been
characterized in the back-end signal processing with the spike sorting
method. Two different spike sorting methods are proposed depending on
different recording bandwidth constraints.

#e1#

#s1#
#abstract1823.txt#


Real-time graphics applications require memory organizations featuring
parallel pixel access and low-cost implementation. This work bases on a
nonlinear skew mapping scheme and exploits the correlation between
consecutive requests for pixels to design an efficient parallel memory
organization. The mapping achieves parallel access, of mn pixels in
various shapes, to the memory organized with mn banks. The proposed
design technique combines the mapping properties and the spatial
correlations among pixel requests to eliminate conflicts by spending at
most one extra cycle every mn consecutive parallel pixel accesses.
Consequently, the technique ensures that any pixel pattern-among these
commonly used in graphics-can be accessed in a single cycle from any
image location. The address computations become straightforward as the
numbers of the requested pixels and the banks-apart from equal-can be
powers of 2.

#e1#

#s1#
#abstract1824.txt#


The inherent limitations of embedded systems make them particularly
vulnerable to attacks. We have developed a hardware monitor that
operates in parallel to an embedded processor and detects any attack
that causes the embedded processor to deviate from its originally
programmed behavior. We explore several different characteristics that
can be used for monitoring and quantify trade-offs between these
approaches. Our results show that our proposed hash-based monitoring
pattern can detect attacks within one instruction cycle at lower memory
requirements than traditional approaches that use control flow
information.

#e1#

#s1#
#abstract1825.txt#


While many-core accelerator architectures, such as today's Graphics
Processing Units (GPUs), offer orders of magnitude more raw computing
power than contemporary CPUs, their massive parallelism often produces
complex dynamic behaviors even with the simplest applications. Using a
fixed set of hardware or simulator performance counters to quantify
behavior over a large interval of time such as an entire application
execution run or program phase may not capture this behavior. Software
and/or hardware designers may consequently miss out on opportunities to
optimize for better performance. Similarly, significant effort may be
expended to find metrics that explain anomalous behavior in
architecture design studies. Moreover, the increasing complexity of
applications developed for today's GPU has created additional
difficulties for software developers when attempting to identify
bottlenecks of an application for optimization. This paper presents a
novel GPU performance visualization tool, AerialVision, to address
these two problems. It interfaces with the GPGPU-Sim simulator to
capture and visualize the dynamic behavior of a GPU architecture
throughout an application run. Similar to existing performance analysis
tools for CPUs, it can annotate individual lines of source code with
performance statistics to simplify the bottleneck identification
process. To provide further insight, AerialVision introduces a novel
methodology to relate pathological dynamic architectural behaviors
resulting in performance loss with the part of the source code that is
responsible. By rapidly providing insight into complex dynamic
behavior, AerialVision enables research on improving many-core
accelerator architectures and will help ensure applications written for
these architectures reach their full performance potential.

#e1#

#s1#
#abstract1826.txt#


With the increasing amounts of parallel programs, produced to
effectively use the parallel computing resources on CMP(chip
multi-processor), it is becoming more important to study these
programs' data access characteristics on CMP. In this paper, the cache
behaviors of parallel programs running on CMP are simulated and
analyzed . Based on the the great degree of shared data accesses among
different threads and the locality of the shared accesses, an actively
pushing cache strategy based on the awareness of shared data is
proposed. The shared data is actively pushed to L1 caches of the slower
threads before it is needed. This strategy takes the advantage of the
potential high on-chip bandwidth while avoids the increased off-chip
bandwidth demand of the prefetch technique due to the inaccuracy.
Experimental results show that this strategy enhances the processing
speeds of the slower threads. The improvement in performance of the
parallel programs can be up to 15.1%.

#e1#

#s1#
#abstract1827.txt#


For the applications that require real-time processing of high-volume
data streams, the scheduling strategy must be adaptability. The Chain
algorithm focuses solely on minimizing the maximum run-time memory
usage, ignoring the important aspect of output latency. Our aim is to
design a scheduling strategy that minimizes the maximum run-time system
memory, while maintaining the output latency within specified bounds.

#e1#

#s1#
#abstract1828.txt#


The communication world still looking forward for a highly
sophisticated, dynamic, probabilistic, key-independent unique secure
data transfer procedure for an ad hoc network. The job is not at all
easy as it require to synchronise or rather globalise several network
criteria of different vendors producing diversified hardware
specification and complexity. However it has been tried to get above
the limitations and define a highly dependent algorithm which is no
doubt processor-hungry and power-consuming but the advantage is its
ability to reduce the burden of increasing bits to keep away the scope
of cryptanalysis, which is the governing term for sniffing and hacking
a network system. Communication networks now a days demand more and
more security due to large number of network filtration and noise. With
the advent of new type of communication networks like sensor networks
and tough competition of bandwidth consumption it has become necessary
to encrypt the data so that during reception the data can be properly
processed and prevent others to interfere and also restrict itself to
get entangled to malfunctions. Also the availability of enormous
processing speed the semiconductor world has paved the way for a
greater opportunity for parallel computing and implementation of
complex architecture.

#e1#

#s1#
#abstract1829.txt#


This paper analyzes algorithmic characteristics of AES
Encryption/Decryption, and proposes a design methodology for AES
algorithm digital hardware circuit through integrating the technologies
of pipeline and parallel connections with dynamic reconfiguration. And
a dynamic reconfiguration circuit model was built to test out the
design methodology. The results of simulation and verification
experiments prove that this model highlights all the advantages of
dynamic reconfiguration technology, considers the application of
parallel connections, increases the throughput, and excels in
processing speed. Thus it illustrates a good and prospect of
application and extension.

#e1#

#s1#
#abstract1830.txt#


In this paper, based on a research project about billet centering
perforation control system in one steelworks, it introduces a punching
control system of automatic detection and centering. The centering
device of the system mainly consists of two symmetrical and parallel
Laser Probing Displacement Sensors, which are used for scanning the
profile of billet. Data are collected through PLC. Then by picture
processing in WINCC, the center of the section of 230~500 mm billet is
found. Therefore, the problems of low efficiency and errors of manual
centering are solved.

#e1#

#s1#
#abstract1831.txt#


The history of speech recognition refers to some decades ago. Speech
recognition performs by using a number of complicated algorithms.
Real-time and rapid execution of these algorithms is very important. In
this paper, a linear systolic architecture is proposed which can
execute speech recognition algorithms based on Hidden Markov Model
(HMM) in parallel and pipeline forms. The proposed architecture is very
regular, consisting of a set of identical and simple processor
elements, which are connected together locally. In order to evaluate
the proposed architecture, it has been designed by using VHDL code and
synthesized on FPGA (Vertix2p) running at 320.546 MHZ.

#e1#

#s1#
#abstract1832.txt#


In this paper, a practical design of fingerprint recognition combined
with smart card verification for two-factor authentication system is
proposed. Our core algorithm processing unit is based on TI's
TMS320VC5510, a high-performance, low-power and fixed-point DSP.
Feature data storage is completed by RF card, which is different from
ordinary solution. When DSP finishes the feature extracting, features
are packaged according to transfer protocol and matched as the 1:1
mapping method from the RF card. 5510 has more than one McBSP but no
asynchronous communication interface; we implement a software way to
simulate UART to communicate with RF card, using the McBSP and DMA.
This operation is effective and cost-efficient for the whole system. In
addition, optimization and parallel transfer feasibility speed up the
performance.

#e1#

#s1#
#abstract1833.txt#


As processing power becomes cheaper and more available by using cluster
of computers, the needs for parallel algorithms, which can harness
these computing potentials, are increasing. Automatic database
normalization is an application of parallel algorithms. Normalization
is the most exercised technique for the analysis of relational
databases. It aims at creating a set of relational tables with minimum
data redundancy that preserve consistency and facilitate correct
insertion, deletion, and modification. While existing sequential
algorithms are usually much time consuming, especially the process of
transforming relations into 3NF, in this paper, we have proposed
parallel algorithms for automatic database normalization. The proposed
algorithms have been examined with MPI and its implementation results
on EDM showed that parallel approach reduces the time, efficiently.
Exploiting p processors has reduced the time of Automatic Database
Normalization to (n/sup 2/ .m)/p + c in which c is the communication
overhead between the processors, m is the number of simple keys, and n
is the number of determinant keys.

#e1#

#s1#
#abstract1834.txt#


A novel approach for counting people passing through a region where is
surveilled by a parallel-fixed binocular camera is proposed in this
paper. Firstly, two virtual counting lines are set to obtain four line
space-time gray images. Then an integrated algorithm based on binocular
stereovision and regional gray projection is proposed to extract and
split overlapped moving objects in complex counting environment. At
last, the object matching and counting is accomplished in space-time
disparity maps. The counting algorithm takes full advantage of the
objects' 2D and 3D information, and is implemented on the DSP hardware
system. The practical test results show that the proposed system can
operate in real-time and achieve a counting accuracy of over 96% in
crowded people environment.

#e1#

#s1#
#abstract1835.txt#


A new parallel programming framework for DNA sequence alignment in
homogeneous multi-core processor architectures is proposed. Contrasting
with traditional coarse-grained parallel approaches, that divide the
considered database in several smaller subsets of complete sequences to
be aligned with the query sequence, the presented methodology is based
on a slicing procedure of both the query and the database sequence
under consideration in several tiles/chunks that are concurrently
processed by the several cores available in the multi-core processor.
The obtained experimental results have proven that significant
accelerations of traditional biological sequence alignment algorithms
can be obtained, reaching a speedup that is linear with the number of
available processing cores and very close to the theoretical maximum.

#e1#

#s1#
#abstract1836.txt#


In grid and parallel systems, backfilling has proven to be a very
efficient method when parallel jobs are considered. However, it
requires a job's runtime to be known in advance, which is not realistic
with current technology. Various prediction methods do exist, but none
is very accurate. In this study, we examine a grid system where both
parallel and sequential jobs require service. Backfilling is used, but
an error margin is added to a job's runtime prediction. The impact on
system performance is examined and the results are compared with the
optimal case of runtimes being known. Two different scheduling
techniques are considered and a simulation model is used to evaluate
system performance.

#e1#

#s1#
#abstract1837.txt#


The following topics are dealt with: KNN queries; distributed data;
stream mining; location based services; probabilistic databases;
spatial indexing; privacy techniques; skyline queries; information
integration; query interfaces; workflow management; workload
management; indexing; hashing; data mining; database reliability;
spatial databases; sensor networks; query optimization; graph mining;
parallel processing; World Wide Web; collaborative applications; and
social networks.

#e1#

#s1#
#abstract1838.txt#


Massive data analysis on large clusters presents new opportunities and
challenges for query optimization. Data partitioning is crucial to
performance in this environment. However, data repartitioning is a very
expensive operation so minimizing the number of such operations can
yield very significant performance improvements. A query optimizer for
this environment must therefore be able to reason about data
partitioning including its interaction with sorting and grouping. S

#e1#

#s1#
#abstract1839.txt#


We propose a multithreshold progressive reconstruction method. The
image is encoded three times using Joint Photographic Experts Group
(JPEG): first with a low-quality factor, then with a medium-quality
factor, and last with a high-quality factor. Huffman coding is employed
to encode the difference between the important image and the
high-quality JPEG decompressed image. The three JPEG codes and the
Huffman code are shared, respectively, according to four prespecified
thresholds. The n-generated equally important shadows can be stored or
transmitted using n channels in parallel. Cooperation among these
generated shadows can progressively reconstruct the important image.
The reconstructed image is loss-free when the number of collected
shadows reaches the largest threshold. Each shadow is very compact and
so can be hidden successfully in the JPEG codes of cover images to
reduce the probability of being attacked when transmitted in an
unfriendly environment. Comparisons with other image sharing methods
are made. The contributions, such as easiness to apply to scalable
Moving Picture Experts Group (MPEG) video transmission or resistance to
differential attack, are also included.

#e1#

#s1#
#abstract1840.txt#


This paper presents active current-sharing control approaches for
parallel-connected AC-to-DC power system architectures consisting of
multiple power-processing channels, each of which comprises a cascade
connection of a front-end active power factor correction (APFC) stage
and an isolated back-end DC-DC converter. By employing a
current-sharing method to the back-end converters, current-mode
commercial-off-the-shelf (

#e1#

#s1#
#abstract1841.txt#


Distributed applications are often faced with a choice between improved
throughput or improved reliability but not both. We argue that this is
not a strict dichotomy and propose a framework that improves both
application performance and reliability by adaptively adjusting source
and channel coding parameters. For simplicity, we assume that the
source and channel we work with are memoryless and in general behave in
a way such that Shannon's separation theorem holds. Although not all
sources and channels could be characterized as such, doing so allows us
to work this resource allocation problem in parallel. We reduce
redundant transmissions and minimize bandwidth utilization through an
LZ77 style dictionary based source coding approach. To ensure data
integrity, we apply rateless forward error correction techniques at the
transport layer. Our algorithm works in conjunction with physical layer
forward error correction and generates just enough overhead needed to
achieve error free transmission without requiring a heavy use of a
reverse channel for acknowledgments. We show through simulations that
our combined source and channel approach reduces network traffic in our
experimental platform by a measurable amount while maintaining and at
times exceeding the Quality of Service (QoS) that is obtained without
our technique.

#e1#

#s1#
#abstract1842.txt#


Ma and Sonka proposed a fully parallel 3D thinning algorithm which does
not always preserve topology. We propose an algorithm based on P-simple
points which automatically corrects Ma and Sonka's algorithm. As far as
we know, our algorithm is the only fully parallel curve thinning
algorithm which preserves topology.

#e1#

#s1#
#abstract1846.txt#


This work presents a regularization technique applied to an inverse
radiative transfer problem formulated as a finite dimensional
optimization problem and solved by a hybridization of the ant colony
optimization (A

#e1#

#s1#
#abstract1847.txt#


A discrete-continuous problem of non-preemptive task scheduling on
identical parallel processors is considered. Tasks are described by
means of a dynamic model, in which the speed of the task performance
depends on the amount of a single continuously divisible renewable
resource allotted to this task over time. An upper bound on the
completion time of all the tasks is given. The criterion is to minimize
the maximum resource consumption at each time instant, i.e., the
resource level. This problem has been observed in many industrial
applications, where a continuously divisible resource such as gas,
fuel, electric, hydraulic or pneumatic power, etc., has to be
distributed among the processing units over time, and it affects their
productivity. The problem consists of two interrelated subproblems:
task sequencing on processors (discrete subproblem) and resource
allocation among the tasks (continuous subproblem). An optimal resource
allocation algorithm for a given sequence of tasks is presented and
computationally tested. Furthermore, approximation algorithms are
proposed, and their theoretical and experimental worst-case
performances are analyzed. Computer experiments confirmed the
efficiency of all the algorithms. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1866.txt#


Online Analytical Processing (OLAP) has become a primary component of
today's pervasive Decision Support systems. As the underlying databases
grow into the multi-terabyte range, however, single CPU OLAP servers
are being stretched beyond their limits. In this paper, we present a
comprehensive model for a fully parallelized OLAP server. Our
multi-node platform actually consists of a series of largely
independent sibling servers that are "glued" together with a
lightweight MPI-based Parallel Service Interface (PSI). Physically, we
target the commodity-oriented, "shared nothing" Linux cluster, an
architecture that provides an extremely cost effective alternative to
the "shared everything" commercial platforms often used in high-end
database environments. Experimental results demonstrate both the
viability and robustness of the design. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1908.txt#


We consider the parallel interpretation of relaxation method for
algebraic reconstruction of tomographic images and a mathematical model
of method parallel interpretation and its advantages over direct
reconstruction methods.

#e1#

#s1#
#abstract1909.txt#


The LHCb Experiment is a hadronic precision experiment at the LHC
accelerator aimed at mainly studying b-physics by profiting from the
large b-anti-b-production at LHC. The challenge of high trigger
efficiency has driven the choice of a readout architecture allowing the
main event filtering to be performed by a software trigger with access
to all detector information on a processing farm based on commercial
multi-core PCs. The readout architecture therefore features only a
relatively relaxed hardware trigger with a fixed and short latency
accepting events at 1 MHz out of a nominal proton collision rate of 30
MHz, and high bandwidth with event fragment assembly over Gigabit
Ethernet. A fast central system performs the entire synchronization,
event labelling and control of the readout, as well as event management
including destination control, dynamic load balancing of the readout
network and the farm, and handling of special events for calibrations
and luminosity measurements. The event filter farm processes the events
in parallel and reduces the physics event rate to about 2 kHz which are
formatted and written to disk before transfer to the offline
processing. A spy mechanism allows processing and reconstructing a
fraction of the events for online quality checking. In addition a 5 Hz
subset of the events are sent as express stream to offline for checking
calibrations and software before launching the full offline processing
on the main event stream.

#e1#

#s1#
#abstract1910.txt#


This paper introduces the Graphite open-source distributed parallel
multicore simulator infrastructure. Graphite is designed from the
ground up for exploration of future multi-core processors containing
dozens, hundreds, or even thousands of cores. It provides high
performance for fast design space exploration and software development.
Several techniques are used to achieve this including: direct
execution, seamless multicore and multi-machine distribution, and lax
synchronization. Graphite is capable of accelerating simulations by
distributing them across multiple commodity Linux machines. When using
multiple machines, it provides the illusion of a single process with a
single, shared address space, allowing it to run off-the-shelf pthread
applications with no source code modification. Our results demonstrate
that Graphite can simulate target architectures containing over 1000
cores on ten 8-core servers. Performance scales well as more machines
are added with near linear speedup in many cases. Simulation slowdown
is as low as 41* versus native execution.

#e1#

#s1#
#abstract1911.txt#


SIMD devices have gained widespread acceptance in modern microprocessor
designs for their superior performance for multimedia applications.
However, there are three remaining limitations to the efficient
utilization of SIMD devices in general-purpose computer systems: memory
alignment, data reorganization and control flow. This paper presents
SIF, an efficient SIMD interface framework that addresses these three
shortcomings without modifying existing ISA. It is designed around a
permutation vector register file (PVRF) and it adds new extended
instructions to set internal permutation state in SIMD datapath rather
than putting the permutation state setting bits in every instruction.
The implicit permutation capability provided by PVRF results in zero
overhead, which frees the handling of three limitations by using
permutation instructions. To further reduce the state setting
instructions in SIMD datapath, a technique that moves the workloads
from SIMD pipeline into scalar pipeline is also introduced. With the
help of proposed compilation algorithm, SIF can efficiently transform
regular SIMD codes into SIF codes which make it easily integrated in
all existing SIMD devices. We implemented these techniques in a
vectorizing compiler and experimental results show that most of the
permutation overhead instructions can be eliminated and distinct
performance speedup can be achieved, which is 37% higher than current
SIMD techniques on average.

#e1#

#s1#
#abstract1912.txt#


Thread-Level Speculation (TLS) has been proposed to facilitate the
extraction of parallel threads from sequential applications. Most prior
work on TLS has focused on architectural features directly related to
supporting the main TLS operations. In this work we, instead,
investigate how a common microarchitectural feature, namely branch
prediction, interacts with TLS. We show that branch prediction for TLS
is even more important than it is for sequential execution.
Unfortunately, branch prediction for TLS systems is also inherently
harder. Code partitioning and re-executions of squashed threads pollute
the branch history making it harder for predictors to be accurate. We
thus propose to augment the hardware, so as to accommodate Multi-Path
Execution (MP) within the existing TLS protocol. Under the MP execution
model, all paths following a number of hard-to-predict conditional
branches are followed simultaneously. MP execution thus removes
branches that would have been otherwise mispredicted, helping in this
way the core to exploit more ILP. We show that, with only minimal
hardware support, one can combine these two execution models into a
unified one. Experimental results show that our combined execution
model achieves speedups of up to 23.2%, with an average of 9.2%, over
an existing state-of-the-art TLS system and speedups of up to 138 %,
with an average of 28.2%, when compared with MP execution for a subset
of the SPEC2000 Int benchmark suite.

#e1#

#s1#
#abstract1921.txt#


Despite its high significance, the clinical utilization of image
registration remains limited because of its lengthy execution time and
a lack of easy access. The focus of this work was twofold. First, we
accelerated our course-to-fine, volume subdivision-based image
registration algorithm by a novel parallel implementation that
maintains the accuracy of our uniprocessor implementation. Second, we
developed a thin-client computing model with a user-friendly interface
to perform rigid and nonrigid image registration. Our novel parallel
computing model uses the message passing interface model on a 32-core
cluster. The results show that, compared with the uniprocessor
implementation, the parallel implementation of our image registration
algorithm is approximately 5 times faster for rigid image registration
and approximately 9 times faster for nonrigid registration for the
images used. To test the viability of such systems for clinical use, we
developed a thin client in the form of a plug-in in OsiriX, a
well-known open source PACS workstation and DI

#e1#

#s1#
#abstract1922.txt#


Simulations of particles which are emitted in laser ablation have been
performed by the method of Direct Simulation Monte Carlo to investigate
the deposition profiles of the emitted particles. The influences of the
temperature, pressure and stream velocity of the initial evaporated
layer formed during laser ablation process on the profile of the
deposited film have been examined. It is found that the temperature
gives a minor influence on the deposition profile, whereas the stream
velocity and the pressure of the initial evaporated layer have a
greater impact on the deposition profile. The energy in the direction
of surface normal (E/sub perp /) and that in the parallel direction of
the surface (E/sub ||/) are shown to increase and decrease,
respectively after the laser irradiation due to collisions between the
emitted particles, and this trend is magnified as the pressure
increases. As a consequence, the stream velocity in the direction of
surface normal increases with the increase in the pressure. A mechanism
of the phenomenon that a metal with a lower sublimation energy shows a
broader angular distribution of emitted particles is presented. It is
suggested that low density of evaporated layer of a metal with a low
sublimation energy at its melting point decreases the number of
collisions in the layer, leading to the low stream velocity in the
direction of surface normal, which results in the broader deposition
profile of the emitted particles. [All rights reserved Elsevier].

#e1#

#s1#
#abstract1923.txt#


We address the problem of simplifying Portuguese texts at the sentence
level by treating it as a "translation task". We use the Statistical
Machine Translation (SMT) framework to learn how to translate from
complex to simplified sentences. Given a parallel corpus of original
and simplified texts, aligned at the sentence level, we train a
standard SMT system and evaluate the "translations" produced using both
standard SMT metrics like BLEU and manual inspection. Results are
promising according to both evaluations, showing that while the model
is usually overcautious in producing simplifications, the overall
quality of the sentences is not degraded and certain types of
simplification operations, mainly lexical, are appropriately captured.

#e1#

#s1#
#abstract1924.txt#


The last few years have seen many significant progresses in the field
of application-specific processors. One example is network security
processors (NSPs) that perform various cryptographic operations
specified by network security protocols and help to offload the
computation intensive burdens from network processors (NPs). This
article presents a high performance NSP system architecture
implementation intended for both internet protocol security (IPSec) and
secure socket layer (SSL) protocol acceleration, which are widely
employed in virtual private network (VPN) and e-commerce applications.
The efficient dual one-way pipelined data transfer skeleton and
optimised integration scheme of the heterogenous parallel crypto engine
arrays lead to a Gbps rate NSP, which is programmable with domain
specific descriptor-based instructions. The descriptor-based control
flow fragments large data packets and distributes them to the crypto
engine arrays, which fully utilises the parallel computation resources
and improves the overall system data throughput. A prototyping platform
for this NSP design is implemented with a Xilinx XC3S5000 based FPGA
chip set. Results show that the design gives a peak throughput for the
IPSec ESP tunnel mode of 2.85 Gbps with over 2100 full SSL handshakes
per second at a clock rate of 95 MHz.

#e1#

#s1#
#abstract1925.txt#


Set-partitioning in hierarchical trees (SPIHT) is one of the well-known
image compression schemes. SPIHT offers an agreeable compression ratio
and produces an embedded bit-stream for progressive transmission.
However, the major disadvantage of SPIHT is its large memory
requirement. In this paper, we propose a memory efficient SPIHT image
coder and its parallel implantation. The memory requirement is reduced
without sacrificing image quality. All bit-planes are concurrently
encoded in order to speed up the entire coding flow. The result shows
that the proposed algorithm is roughly 6 times faster than the original
SPIHT. For a 512 * 512 image, the memory requirement is reduced from
5.83 Mb to 491 Kb. The proposed algorithm is also realized on FPGA.
With pipeline design, the circuit can run at 110 MHz, which can encode
a 512 * 512 image in 1.438 ms. Thus, the circuit achieves very high
throughput, 182 MPixels/sec, and can be applied to high performance
image compression applications.

#e1#

#s1#
#abstract1926.txt#


The algebraic path problem (APP) is a general framework which unifies
several solution procedures for a number of well-known matrix and graph
problems. In this paper, we present a new 3-dimensional (3D) orbital
algebraic path algorithm and corresponding 2-D toroidal array
processors which solve the n x n APP in the theoretically minimal
number of 3n time-steps. The coordinated time-space scheduling of the
computing and data movement in this 3-D algorithm is based on the
modular function which preserves the main technological advantages of
systolic processing: simplicity, regularity, locality of
communications, pipelining, etc. Our design of the 2-D systolic array
processors is based on a classical 3-D rarr 2-D space transformation.
We have also shown how a data manipulation (copying and alignment) can
be effectively implemented in these array processors in a
massively-parallel fashion by using a matrix-matrix multiply-add
operation.

#e1#

#s1#
#abstract1937.txt#


In parallel query-processing environments, accurate, time-oriented
progress indicators could provide much utility given that inter- and
intra-query execution times can have high variance. However, none of
the techniques used by existing tools or available in the literature
provide non-trivial progress estimation for parallel queries. In this
paper, we introduce Parallax, the first such indicator. While several
parallel data processing systems exist, the work in this paper targets
environments where queries consist of a series of MapReduce jobs.
Parallax builds on recently-developed techniques for estimating the
progress of single-site SQL queries, but focuses on the challenges
related to parallelism and variable execution speeds. We have
implemented our estimator in the Pig system and demonstrate its
performance through experiments with the PigMix benchmark and other
queries running in a real, small-scale cluster.

#e1#

#s1#
#abstract1940.txt#


Hellfire framework (HellfireFW) presents a design flow for the design
of MPSoC based critical embedded systems. Health-care electronics,
security equipment and space aircraft are examples of such systems
that, besides presenting typical embedded system's constraints, bring
new design challenges as their restrictions are even tighter in terms
of area, power consumption and high-performance in distributed
computing involving real-time processing requirement. In this paper, we
present the Hellfire framework, which offers an integrated tool-flow in
which design space exploration (DSE), OS customization and static and
dynamic application mapping are highly automated. The designer can
develop embedded sequential and parallel applications while evaluating
how design decisions impact in overall system behavior, in terms of
static and dynamic task mapping, performance, deadline miss ratio,
communication traffic and energy consumption. Results show that: i) our
solution is suitable for hard real-time critical embedded systems, in
terms of real-time scheduling and OS overhead; ii) an accurate analysis
of critical embedded applications in terms of deadline miss ratio can
be done using HellfireFW; iii) designer can better decide which
architecture is more suitable for the application; iv) different HW/SW
solutions by configuring both the RTOS and the HW platform can be
simulated.

#e1#

#s1#
#abstract1941.txt#


Due to the extremely large sizes of power grids, IR drop analysis has
become a computationally challenging problem both in terms of runtime
and memory usage. In order to design scalable algorithms to handle ever
increasing power-grid sizes, the most promising approach is to use a
"divide-and-conquer" strategy such as domain decomposition. Such an
approach not only decomposes a large problem into manageable
sub-problems, it also naturally allow a parallel processing solution
for further speedup in computation time. As a result, a power-grid
analysis algorithm based upon the traditional domain decomposition
method has been reported. Unfortunately, the method in has strong
limitation on the size of the interfaces between the sub-problems and
therefore severely limits its capability in solving very large
problems. In this paper, we present a block-iterative
domain-decomposition algorithm which effectively combines the
advantages of direct solvers and iterative methods. With a carefully
chosen domain decomposition strategy, our approach does not suffer from
the difficulties. While the algorithm fails to analyze a power grid of
4 millions nodes, our algorithm solves a power grid of 42 millions
nodes accurately in 1.5 hours.

#e1#

#s1#
#abstract1942.txt#


The authors introduce the iterative leap-field domain decomposition
method that is tailored to the finite element method, by combining the
concept of domain decomposition and the Huygens' Principle. In this
method, a large-scale electromagnetic boundary value problem is
partitioned into a number of suitably-defined 'small' and manageable
subproblems whose solutions are assembled to obtain the global
solution. The main idea of the method is the iterative application of
the Huygens' Principle to the fields radiated by the equivalent
currents calculated in each iteration. In the context of the
electromagnetic scattering, the method can be applied to cases
involving multiple objects, as well as to a 'single' challenging object
in a straightforward manner via the locally conformal perfectly matched
layer technique. The most attractive feature of the method is the
considerable reduction in the memory requirements and computation time.
It is observed that convergence is achieved after a few iterations and
computation time may further be reduced via parallel processing
techniques. After developing the analytical background of this method,
we present some numerical results related to the three-dimensional
electromagnetic scattering problems.

#e1#

#s1#
#abstract1943.txt#


In order for the video decoding processing such as H.264 and VC-1 to be
effective in multi-core environments, several kinds of parallelisms
must be utilised. Here, a novel parallelisation methodology, macroblock
row-level parallelism (MBRLP), of video decoding is presented. The ETRI
multimedia processing core (EMC) and the ETRI multi-core platform (EMP)
are proposed for adopting MBRLP. In terms of the scalability and
utilisation of processing cores, MBRLP has advantages over other
parallelisation strategies such as frame, slice and macroblock
(MB)-level parallelism. The scalability can be easily achieved by just
increasing the number of processing cores and applying homogeneous
software design/optimisation techniques to each EMC. Instead of
employing a dynamic MB-level scheduler, a hybrid approach is used,
which is a two-stage functional pipelining combined with MBRLP. The
hybrid approach of combining MBRLP and de-blocking pipelining can
relieve the synchronisation and inter-processor communication overheads
incurred by multicore decoding systems as well as run-time scheduler's
overheads. As a result, the proposed parallelisation method and
architectures can boost the performance with the efficiency of 83%. The
proposed architecture consisting of six EMC clusters has the capability
to process Dl (720 * 480) 30 fps real-time decoding at around 200 MHz.
The same concept can be applied to full-HD (1920 * 1088) video decoding
in this work. It can be found that as the number of processing cores
increase, the performance improvement is enhanced almost linearly. The
EMP consisting of four EMC clusters (eight cores), memories and other
peripherals are prototyped on Xilinx Virtex4 XC4VL200 FPGA which is
operating at 60 MHz.

#e1#

#s1#
#abstract1944.txt#


This paper presents a system model and method for the 2-D imaging
application via a narrowband multiple-input multiple-output (MIMO)
radar system with two perpendicular linear arrays. Furthermore, the
imaging formulation for our method is developed through a Fourier
integral processing, and the parameters of antenna array including the
cross-range resolution, required size, and sampling interval are also
examined. Different from the spatial sequential procedure sampling the
scattered echoes during multiple snapshot illuminations in inverse
synthetic aperture radar (ISAR) imaging, the proposed method utilizes a
spatial parallel procedure to sample the scattered echoes during a
single snapshot illumination. Consequently, the complex motion
compensation in ISAR imaging can be avoided. Moreover, in our array
configuration, multiple narrowband spectrum-shared waveforms coded with
orthogonal polyphase sequences are employed. The mainlobes of the
compressed echoes from the different filter band could be located in
the same range bin, and thus, the range alignment in classical ISAR
imaging is not necessary. Numerical simulations based on synthetic data
are provided for testing our proposed method.

#e1#

#s1#
#abstract1945.txt#


Purpose - With the availability of powerful personal computers (PCs),
workstations and networking devices, the recent trend in parallel
computing is to connect a number of individual workstations (PC and PC
symmetric multiprocessor systems (SMP)) to solve computation-intensive
tasks in parallel way on such clusters (networks of workstations (NOW),
SMP and Grid). In this sense, it is not more true to consider
traditionally evolved parallel computing and distributed computing as
two separate research disciplines. Current trends in high performance
computing are to use NOW (and SMP) as a cheaper alternative to
traditionally used massively parallel multiprocessors or supercomputers
and to profit from unifying of both mentioned disciplines. The purpose
of this paper is to consider the individual workstations could be so
single PC as parallel computers based on modern SMP implemented within
workstation. Design/methodology/approach - Such parallel systems (NOW
and SMP), are connected through widely used communication standard
networks and co-operate to solve one large problem. Each workstation is
threatened similarly to a processing element as in a conventional
multiprocessor system. But, personal processors or multiprocessors as
workstations are far more powerful and flexible than the processing
elements in conventional multiprocessors. To make the whole system
appear to the applications as a single parallel computing engine (a
virtual parallel system), run-time environments such as OpenMP, Java
(SMP), message passing interface, Java (NOW) are used to provide an
extra layer of abstraction. Findings - To exploit the parallel
processing capability of such cluster, the application program must be
paralleled. The effective way how to do it for (parallelisation
strategy) belongs to a most important step in developing effective
parallel algorithm (optimisation). To behaviour analysis, all overheads
that have the influence to performance of parallel algorithms
(architecture, computation, communication, etc.) have to be taken into
account. In this paper, such complex performance evaluation of
iterative parallel algorithms (IPA) and their practical implementations
are discussed (Jacobi and Gauss-Seidel iteration). On real application
example, the various influences in process of modelling and performance
evaluation and the consequences of their distributed parallel
implementations are demonstrated. Originality/value - The paper
usefully shows that better load balancing can be achieved among used
network nodes (performance optimisation of parallel algorithm).
Generally, it claims that the parallel algorithms or their parts
(processes) with more communication (similar to analyzed Gauss-Seidel
parallel algorithm) will have better speed-up values using modern SMP
parallel system as its parallel implementation in NOW. For the
algorithms or processes with small communication overheads (similar to
analysed Jacobi parallel algorithm) the other network nodes can be used
based on single processors.

#e1#

#s1#
#abstract1946.txt#


We report on the effects of CF/sub 4/ plasma process in technical RF
barrel reactor on surface termination and the resulting electronic
surface barrier of boron-doped diamond in electrolytes. The surface
characteristics were evaluated for epitaxial single crystalline layers
with sub-nm roughness. The capacitance-voltage characteristics of the
processes electrodes implied a low electronic barrier at the
fluorine-terminated areas, comparable to hydrogen termination in
literature. However carbon-oxygen groups were still present on approx.
15% of the surface area after the plasma process. To analyse the
electronic barrier of the fluorinated diamond, we proposed an
electrical model of two parallel metal-oxide-semiconductor structures.
This model included fixed parameters derived from the analysis of a
diamond electrode exposed to plasma oxidation. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract1947.txt#


The interaction between active observer and environment is connected
with the ambiguity of visual signals interpretation. The ambiguity can
be treated as action motivation. Perspective picture variations on
sensitive visual matrix can be concerned with dynamic environment or
visual system movements. To define variation components, corresponded
to dynamic environmental processes, observer need to have internal
environment model and to forecast perspective picture variations while
movement. If the forecast does not correspond to the observed
perspective picture, the environment is dynamic. Another type of
ambiguity appears in recognition tasks. Some objects have the same
perspective projections and have to be observed from different points
of view to be recognised. Therefore, recognition process is treated as
the sequence of procedures: forecast, verification, decision.
Three-dimensional objects representation is possible with active
monocular sensor or with binocular visual system. In any case, the
correspondent point problem is to be solved to make three-dimensional
objects representation. Correspondent points are to be found on the
perspective pictures sequence. Three-dimensional point coordinates (in
internal coordinate system) can be defined if two perspective
correspondent points are found and observer movement vector is known.
Observer movements set restriction rules to forecast movements of
correspondent points on perspective projections. Correspondent points
have to be searched on parallel lines if visual system configuration is
binocular and have to be searched on radial rays if monocular visual
system moves in view direction. Forecast is modeled as visual system
presetting. If presetting procedures are involved in cognitive cycle,
the dynamic environment perception can be performed.

#e1#

#s1#
#abstract1948.txt#


Predicting protein mutant stability changes is important for protein
design. Although many methods have tried to improve prediction accuracy
by various models, it will be difficult to employ them when the
required input information is incomplete. Therefore, we have proposed a
fuzzy query method previously to predict stability changes upon single
mutations using partial input information. However, when a researcher
has tens or hundreds of queries to predict, he needs to input query
data repeatedly. It is better to issue queries in a batch, which incurs
another problem. Because our proposed fuzzy query method is implemented
in the FuzzyCLIPS language, it takes a long time to process a batch of
queries. To shorten the response time, we propose parallelizing queries
in emerging cluster systems or grid systems. Furthermore, to address
the problem of heterogeneous computing power of cluster and grid nodes,
a variety of self-scheduling schemes have been implemented for better
load balance. Experimental results show that our proposed parallel
fuzzy query system can provide hundreds of predications in more
reasonable response time.

#e1#

#s1#
#abstract1949.txt#


This paper has put forward a type of parallel rendering algorithm for
large-scale terrain. Based on Sort-First tactics, this paper has
proposed a load balancing algorithm to achieve relatively small extra
overhead and relatively balanced load distribution while guarantee
relatively high load calculation and distribution efficiency at the
same time. The modified nested grid algorithm for terrain rendering
decreases the pressure on calculating view frustum culling during load
balancing while increases rendering efficiency at the same time. In
this paper, we use JPEG lossy compression to process image data to be
transmitted, and use effective synchronization mechanism between two
nodes, so as to further improve system efficiency.

#e1#

#s1#
#abstract1950.txt#


Aligned parallel corpora are an important resource for a wide range of
multilingual researches, specifically, corpus-based machine
translation. In this paper we present a Persian-English
sentence-aligned parallel corpus by mining Wikipedia. We propose a
method of extracting sentence-level alignment by using an extended
link-based bilingual lexicon method. Experimental results show that our
method increase precision, while it reduce the total number of
generated candidate pairs.

#e1#

#s1#
#abstract1951.txt#


This paper proposes a methodology based on formal correspondence
checking to automatically debug and also optimize pipelined
microprocessors including reconfigurable processors with timing error
recovery techniques. Since formal verification analyzes the design
exhaustively, it may give good insights into not only debugging but
also optimization of hardware designs with complicated control
structures. The paper gives two main contributions, (1) modeling and
formal verification of pipelined microprocessors including
reconfigurable processors with timing error recovery techniques and (2)
an approach to debug and optimize the implementation using the UCLID
system as a correspondence checker. Using our method, the debug time
can be reduced significantly. In addition, the implementation can be
optimized by removing unnecessary signals and components while the
correctness of the design is guaranteed.

#e1#

#s1#
#abstract1952.txt#


Coarse Grained Reconfigurable Array (CGRA) architectures give high
throughput and data reuse for regular algorithms while providing
flexibility to execute multiple algorithms on the same architecture.
This paper investigates systolic mapping techniques for mapping
biosignal processing algorithms to CGRA architectures. A novel
methodology using synchronous data flow (SDF) graphs and control and
data flow (CDF) graphs for mapping is presented. Mapping signal
processing algorithms in this manner is shown to give up to a 88%
reduction in memory accesses and significant savings in fetch and
decode operations while providing high throughput.

#e1#

#s1#
#abstract1953.txt#


High-Performance Reconfigurable Computers (HPRCs) are parallel machines
consisting of FPGAs and microprocessors, with the FPGAs used as
co-processors. The execution of parallel applications on such systems
has mainly followed the Single-Program Multiple-Data (SPMD) model;
however, overall system resources are often underutilized because of
the asymmetric distribution of the reconfigurable (co-)processors
relative to the (main) processors. Furthermore, with the introduction
of HPRCs containing multi/many-core technologies, underutilization of
system resources becomes more obvious especially for multi-tasking and
multi-user usage. To address the asymmetry problem, we propose a
resource virtualization solution based on Partial Run-Time
Reconfiguration (PRTR). The proposed technique allows space, time,
and/or spacetime sharing of the reconfigurable (co-)processors among
the (main) processors and thus increasing the overall system
utilization. We show the effectiveness of the proposed concepts through
a stochastic execution model verified with experimental implementations
on the Cray XD1 platform. The results demonstrate favorable performance
as well as scalability characteristics.

#e1#

#s1#
#abstract1954.txt#


The SIMD parallel systems play a crucial role in the field of intensive
signal processing. For most the parallel systems, communication
networks are considered as one of the challenges facing researchers.
This work describes the FPGA implementation of two reconfigurable and
flexible communication networks integrated into mppSoC. An mppSoC
system is an SIMD massively parallel processing System on Chip designed
for data-parallel applications. Its most distinguished features are its
parameterization and the reconfigurability of its interconnection
networks. This reconfigurability allows to establish one configuration
with a network topology well mapped to the algorithm communication
graph so that higher efficiency can be achieved. Experimental results
for mppSoC with different communication configurations demonstrate the
performance of the used reconfigurable networks and the effectiveness
of algorithm mapping through reconfiguration.

#e1#

#s1#
#abstract1955.txt#


Our work aims at adapting the concept of virtualization, which is known
from the context of operating systems, for concurrent hardware design.
By contrast, the proposed concept applies virtualization not to
processors or applications but to smaller processing units within a
parallel array of homogeneous instances and individual tasks. Thereby,
virtualization during runtime enables fault tolerance without the need
for spare redundancy: The proposed architecture is able to switch
seamlessly between parallelism for execution acceleration and
redundancy for fault tolerance. In addition, faulty instances are
completely decoupled from the system. This allows for an easy dynamic
and partial reconfiguration. Using this concept, self-healing
mechanisms can be implemented, as decoupled, faulty instances may be
replaced by operational instances during reconfiguration. We present
this hardware-based virtualization concept on the basis of a parallel
array of multipliers used for ECC point-multiplications.

#e1#

#s1#
#abstract1956.txt#


This paper presents and compares different load balancing strategies in
multi-core network processor (NP) chips. In our FlexPath NP system,
packets are differentiated according to application-dependent
processing requirements and optimized processing paths are provisioned
for these applications. We derive a novel load balancing mechanism
(S&H) by combining two schemes for stateful and stateless network
applications in order to achieve better overall system throughput and
reduced packet latencies. We show that appropriate QoS for the
different regarded application types can be achieved under varying NP
load conditions, while maintaining an almost uniform utilization of the
available processing resources. Even though the investigations are
focused on the FlexPath NP architecture, the concepts can also be
applied to other architectures, where the incoming load has to be
distributed among several parallel entities within an NP.

#e1#

#s1#
#abstract1957.txt#


Designing embedded systems the way we did 20 years ago is still alive
and well. As expected, with declining costs, embedded systems are
appearing in more and more applications. Advances to the state of the
art in creating such systems, where memory and processor are precious
resources, have continued, and this work is to be applauded. But just
as interestingly, embedded systems are taking on not only problems that
require loosening the memory and processor constraints but also those
problems that push us into the domain of datacenter design. Datacenter
design is where system thinking about heterogeneous parallel
processing, data archiving, mission-critical connectivity, and energy
management are first-order concerns. In this talk I will illustrate the
latter case with some recent explorations and give some attention to
how datacenter thinking can be effectively ushered into this new domain.

#e1#

#s1#
#abstract1958.txt#


This paper first presents a new architecture of MD5, which achieved the
theoretical upper bound on throughput in the iterative architecture.
And then based on the general proposed architecture, this paper
implemented other two different kinds of pipelined architectures which
are based on the iterative technique and the loop unrolling technique
respectively. The latter with 32-stage pipelining reached a throughput
up to 32.035 Gbps on an Altera Stratix II GX EP2SGX90FF FPGA, and the
speedup achieved 194* over the Intel Pentium 4 3.0 processor. At least
to the authors' knowledge, this is the fastest published FPGA-based
design at the time of writing. At last the proposed designs are
compared with other published MD5 designs, the designs in this paper
have obvious advantages both in speed and logic requirements.

#e1#

#s1#
#abstract1959.txt#


Since reducing the diameter is likely to improve the performance of an
interconnection network, the problem of designing interconnection
network with low diameter is still a current research topic. Another
important issue in the design of interconnection networks for massively
parallel computers is scalability. A new hierarchical interconnection
network topology, called Rectangular Twisted Torus Meshes (RTTM
network), is proposed. At the lowest level of RTTM network, the Level-1
sub-network, also called a Basic Module, consists of a mesh connection
of 2/sup m/ * 2/sup m/ nodes. Successively higher level networks are
built by recursively interconnecting a*2a next lower level sub-networks
in the form of a Rectangular Twisted Torus. An appealing property of
the RTTM network is its smaller diameter and shorter average distance,
which implies a reduction in communication delays. The RTTM network
allows the exploitation of computational locality as well as easy
expansion up to a million processors.

#e1#

#s1#
#abstract1960.txt#


The paper presents a multi-CPU implementation of the preconditioned
conjugate gradient algorithm with an algebraic multigrid preconditioner
(PCG-AMG) for an elliptic model problem on a 3D unstructured grid. An
efficient parallel sparse matrix-vector multiplication scheme
underlying the PCG-AMG algorithm is presented for the many-core GPU
architecture. A performance comparison of the parallel solver shows
that a singe Nvidia Tesla C1060 GPU board delivers the performance of a
sixteen node Infiniband cluster and a multi-GPU configuration with
eight CPUs is about 100 times faster than a typical server CPU core.

#e1#

#s1#
#abstract1961.txt#


The following topics are dealt with: data driven application system;
hardware chips; parallel algebraic multigrid solver; graphic processing
unit; time-domain Maxwell's equation; virtual machine performance
evaluation; grid computing; parallel system performance modeling;
finite element analysis; interactive visualization; 3D geological
modeling; fast ARFTIS reconstruction algorithm; parallel programming;
cellular automata; GPU platform; visual software development; compilers.

#e1#

#s1#
#abstract1962.txt#


Graphics Processing Units (GPUs) offer tremendous computational power.
CUDA (Compute Unified Device Architecture) provides a multi-threaded
parallel programming model, facilitating high performance
implementations of general-purpose computations. However, the
explicitly managed memory hierarchy and multi-level parallel view make
manual development of high-performance CUDA code rather complicated.
Hence the automatic transformation of sequential input programs into
efficient parallel CUDA programs is of considerable interest. This
paper describes an automatic code transformation system that generates
parallel CUDA code from input sequential C code, for regular (affine)
programs. Using and adapting publicly available tools that have made
polyhedral compiler optimization practically effective, we develop a
C-to-CUDA transformation system that generates two-level parallel CUDA
code that is optimized for efficient data access. The performance of
automatically generated code is compared with manually optimized CUDA
code for a number of benchmarks. The performance of the automatically
generated CUDA code is quite close to hand-optimized CUDA code and
considerably better than the benchmarks' performance on a multicore CPU.

#e1#

#s1#
#abstract1963.txt#


The novel apparatus described here was developed to investigate the
thermo-mechanical behavior of metallic films on a substrate by
acquiring the wafer curvature. It comprises an optical module producing
and measuring an array of parallel laser beams, a high resolution
scanning stage, a rapid thermal processing (RTP) chamber and several
accessorial gas control modules. Unlike most traditional systems which
only calculate the average wafer curvature, this system has the
capability to measure the curvature locally in 30 ms. Consequently, the
real-time development of biaxial stress involved in thin films can be
fully captured during any thermal treatments such as temperature
cycling or annealing processes. In addition, the multiple parallel
laser beam technique cancels electrical, vibrational and other random
noise sources that would otherwise make an in situ measurement very
difficult. Furthermore, other advanced features such as the in situ
acid treatment and active cooling extend the experimental conditions to
provide new insights into thin film properties and material behavior.

#e1#

#s1#
#abstract1964.txt#


This paper demonstrates a novel method to perform the delivery and
assembly of standard 01005 format (0.016"*0.008", 0.4 mm*0.2 mm)
thin-film resistors and monolithic ceramic capacitors with a
programmable batch assembly process that leads to 100% yield within
tens of seconds, that is high volume manufacturing compatible. By
characterizing the electrical performance of functional
industrial-grade components assembled onto test substrates, we extend
previous stochastic assembly work performed with dummy parts,
validating our assembly methodology.

#e1#

#s1#
#abstract1965.txt#


An attempt is made to improve the accuracy of a multi-channel parallel
acousto-optical spectrum analysis through involving an additional
nonlinear conversion into data processing. For this purpose we
investigate the potential of exploiting a co-directional collinear wave
heterodyning within an analysis of ultra-high-frequency radio-wave
signals. The wave heterodyning under the proposal is performed via
mixing the longitudinal elastic waves of finite amplitudes. It leads to
a two-cascade processing in the single-crystalline cell and makes it
possible to improve the relative frequency resolution of the
acousto-optical spectrum analysis by an order of magnitude at the same
frequency range. Both the theoretical findings and the corresponding
estimations were used in proof-of-principle experiments directed at
creating a new type of acousto-optical cell. The general concept of the
proposed method and the basic conclusions are confirmed by experiments
with the developed acousto-optic cell made of a lead molybdate crystal.

#e1#

#s1#
#abstract1975.txt#


Magnetic resonance imaging (MRI) and spectroscopy (MRS) have
contributed considerably to our understanding of the structure and
function of the human brain, and are essential in modern radiological
practice. The development and applications of MRI systems at fields of
7 Tesla and above provide many new opportunities and technical
challenges. The increased signal strength at higher fields may be used
to obtain higher resolution images, faster images and/or images with
greater contrast to noise . These can be used to improve detection of
lesions, for more accurate assessment of the structural anatomy of the
brain (including higher resolution tractography of white matter), and
improved sensitivity in functional MRI. However, high field MR imaging
is also affected by macroscopic field variations caused by
inhomogeneities of magnetic susceptibility within the body, and these
can degrade spectra and introduce image distortions. Moreover, the
performance of radiofrequency (RF) coils is also affected at higher
fields, and it is more difficult to create uniform RF fields within
large objects. These challenges are being met using various technical
innovations such as dynamic shimming, the use of parallel arrays of
coils, novel spectral-spatial excitation methods, novel pulse sequences
and post-acquisition digital processing. In combination these efforts
promise to allow ultra-high field imaging and spectroscopy to
contribute significantly to our understanding of brain structure and
function in various conditions.

#e1#

#s1#
#abstract1976.txt#


In color flow imaging for medical diagnosis, the inherent trade-off
between frame rate and image quality may often lead to suboptimal
images. Parallel receive beamforming is used to help overcome this
problem, but this introduces artifacts in the images. In addition to
the parallel beamforming artifacts found in B-mode imaging, we have
found that a difference in curvature of transmit and receive beams
gives a bias in the Doppler velocity estimates. This bias causes a
discontinuity in the velocity estimates in color flow images. In this
work, we have shown that interpolation of the autocorrelation estimates
obtained from overlapping receive beams can reduce these artifacts
significantly. Because the autocorrelation function varies quite
slowly, the beams can be acquired with a considerable time difference,
for instance across interleaving groups or across scan planes in a 3-D
scan. We have shown that a high frame rate of color flow images can be
maintained with parallel beam acquisition with minimal deterioration of
the image quality.

#e1#

#s1#
#abstract1977.txt#


A multiband frequency-swept laser source is proposed using wavelength
division multiplexing based on the periodic transfer function of a
scanning Fabry-Perot interferometer. The optical spectra from different
wavelength bands can be combined to create an ultra wideband light
source which may significantly improve the spatial resolution of
optical coherent tomography (OCT). Parallel processing for data
measured in different wavelength bands was used to speed up signal
analysis. A proof-of-concept experiment was conducted to demonstrate
the feasibility of the proposed technique.

#e1#

#s1#
#abstract2015.txt#


Summary form only given: The HyVM project is developing system support
for future heterogeneous chip multiprocessors. Such hybrid hardware
platforms offer opportunities in terms of improved power/performance
properties, but pose challenges to systems technologies due to
heterogeneous processing cores, non-uniform memory access, and complex
software stacks. The HyVM project is creating new hypervisor- and
system-level abstractions in support of providing a uniform program
execution model for future hybrid computing platforms. Rather than
treating accelerators as external devices, the model anticipates future
integrated systems by providing sets of virtual processing units for
use by both accelerator and commodity programs, offering the resource
management support needed to efficiently execute such parallel
multi-core applications, and supplying the tool chains needed, at
hypervisor level, to permit applications to freely use arbitrary
combinations of accelerator and commodity cores. The talk will overview
the HyVM project, review results that range from efficient methods for
virtualizing accelerators, to online techniques for managing
heterogenous system resources, to JIT binary translation for dealing
with diverse accelerator targets. The effort is driven by both
commercial and high performance applications targeting future hybrid
machines.

#e1#

#s1#
#abstract2016.txt#


We present two methods for estimating replacement probabilities without
using parallel corpora. The first method proposed exploits the possible
translation probabilities latent in Machine Readable Dictionaries
(MRD). The second method is more robust, and exploits context
similarity-based techniques in order to estimate word translation
probabilities using the Internet as a bilingual comparable corpus. The
experiments show a statistically significant improvement over non
weighted structured queries in terms of MAP by using the replacement
probabilities obtained with the proposed methods. The context
similarity-based method is the one that yields the most significant
improvement.

#e1#

#s1#
#abstract2017.txt#


Chip multiprocessors (CMPs) are promising candidates for the next
generation computing platforms to utilize large numbers of gates and
reduce the effects of high interconnect delays. One of the key
challenges in CMP design is to balance out the often-conflicting
demands. Specifically, for today's image/video applications and
systems, power consumption, memory space occupancy, area cost, and
reliability are as important as performance. Therefore, a compilation
framework for CMPs should consider multiple factors during the
optimization process. Motivated by this observation, this paper
addresses the energy-aware reliability support for the CMP
architectures, targeting in particular at array-intensive image/video
applications. There are two main goals behind our compiler approach.
First, we want to minimize the energy wasted in executing replicas when
there is no error during execution (which should be the most frequent
case in practice). Second, we want to minimize the time to recover
(through the replicas) from an error when it occurs. This approach has
been implemented and tested using four parallel array-based
applications from the image/video processing domain. Our experimental
evaluation indicates that the proposed approach saves significant
energy over the case when all the replicas are run under the highest
voltage/frequency level, without sacrificing any reliability over the
latter. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2018.txt#


Recently proposed modern technique of a precise spectrum analysis
within an algorithm of the collinear wave heterodyning implies a
two-stage integrated processing, namely, the wave heterodyning of a
signal in a square-law nonlinear medium and then the optical processing
in the same cell. Technical advantage of this approach is in providing
a direct processing of ultra-high-frequency radio-wave signals with
essentially improved frequency resolution. This algorithm can be
realized on a basis of various physical principles, and we consider an
opportunity of involving the potentials of modern acousto-optics for
these purposes. From this viewpoint, one needs a large-aperture
effective acousto-optical cell, which operates in the Bragg regime and
performs the ultra-high-frequency co-directional collinear acoustic
wave heterodyning. The technique under consideration imposes specific
requirements on the cell's material, namely, a high optical quality of
large-size crystalline boules, high-efficient acousto-optical and
acoustic interactions, and low group velocity of acoustic waves
together with square-low dispersive acoustic losses. We focus our
attention on the solid solutions of thallium chalcogenides and take the
TlBr-TlI (thallium bromine - thallium iodine) solution, which forms
KRS-5 cubic-symmetry crystals with the mass-ratio 58% of TlBr to 42% of
TlI. Analysis shows that the acousto-optical cell made of a KRS-5
crystal oriented along the [111] -axis and the corresponding
longitudinal elastic mode for producing the dynamic diffractive grating
in that crystal can be exploited. With the acoustic velocity of about
1.92 mm/ mu s and attenuation of approximately 10 dB/(cm GHz/sup 2/),
similar cell is capable to provide an optical aperture of 50 mm and one
of the highest figures of acousto-optical merit in solid states in the
visible range. Such a cell is rather desirable for applications to
direct parallel multi-channel optical spectrum analysis with
substantially improved frequency resolution.

#e1#

#s1#
#abstract2019.txt#


Background: Quantification of different types of cells is often needed
for analysis of histological images. In our project, we compute the
relative number of proliferating hepatocytes for the evaluation of the
regeneration process after partial hepatectomy in normal rat livers.
Results: Our presented automatic approach for hepatocyte (HC)
quantification is suitable for the analysis of an entire digitized
histological section given in form of a series of images. It is the
main part of an automatic hepatocyte quantification tool that allows
for the computation of the ratio between the number of proliferating
HC-nuclei and the total number of all HC-nuclei for a series of images
in one processing run. The processing pipeline allows us to obtain
desired and valuable results for a wide range of images with different
properties without additional parameter adjustment. Comparing the
obtained segmentation results with a manually retrieved segmentation
mask which is considered to be the ground truth, we achieve results
with sensitivity above 90% and false positive fraction below 15%.
Conclusions: The proposed automatic procedure gives results with high
sensitivity and low false positive fraction and can be applied to
process entire stained sections.

#e1#

#s1#
#abstract2020.txt#


The recent advances in computer animation, motion image processing,
robotics and so on prompted us to analyze computational complexity of
four-dimensional pattern processing. Thus, the research of
four-dimensional automata as a computational model of four-dimensional
pattern processing has also been meaningful. From this viewpoint, we
introduced a four-dimensional alternating Turing machine (4-ATM)
operating in parallel. In this paper, we continue the investigations
about 4-ATM's, deal with a four-dimensional synchronized alternating
Turing machine (4-SATM), and investigate some properties of 4-SATM's
which each sidelength of each input tape is equivalent. The main topics
of this paper are: (1) hierarchies based on the number of processes of
4-SATM's, and (2) recognizability of connected pictures by 4-SATM's.

#e1#

#s1#
#abstract2021.txt#


In vivo retinal imaging is an outstanding tool to observe biological
processes unfold in real-time. The ability to image microstructure in
vivo can greatly enhance our understanding of function in retinal
microanatomy under normal conditions and in disease. Transgenic mice
are frequently used for mouse models of retinal diseases. However,
commercially available retinal imaging instruments lack the optical
resolution and spectral flexibility necessary to visualize detail
comprehensively. We developed an adaptive optics scanning laser
ophthalmoscope (AO-SLO) specifically for mouse eyes. Our SLO is a
sensor-less adaptive optics system (no Shack Hartmann sensor) that
employs a stochastic parallel gradient descent algorithm to modulate a
deformable mirror, ultimately aiming to correct wavefront aberrations
by optimizing confocal image sharpness. The resulting resolution allows
detailed observation of retinal microstructure. The AO-SLO can resolve
retinal microglia and their moving processes, demonstrating that
microglia processes are highly motile, constantly probing their
immediate environment. Similarly, retinal ganglion cells are imaged
along with their axons and sprouting dendrites. Retinal blood vessels
are imaged both using evans blue fluorescence and backscattering
contrast.

#e1#

#s1#
#abstract2022.txt#


A novel optical scanning method for an anterior segment optical
coherence tomography (AS-OCT) system has been described. This method
has been designed for imaging the entire anterior segment of the eye
(from cornea to posterior surface of the crystalline lens) at a time.
The ability to image the entire anterior segment is crucial in
understanding the mechanism of human accommodation and the efficacy of
accommodative intraocular lenses. In a conventional scanning system the
beam is shined straight into the eye parallel to the optical axis. For
anterior segment imaging, large lateral scan area leads to an increase
in the angle of incidence on each of the four ocular surfaces. This
causes significant reduction in signal reflected from regions further
from the optical axis. This reduction combined with loss in signal due
to coherence makes it difficult to image the entire anterior segment of
the eye, where optical depth penetration of 10mm is required. To
overcome this limitation, we have designed a new OCT scanning system,
which achieves close to normal incidences across all the lateral
locations on the ocular surfaces within a 6 mm clear aperture. This
provides an increase in the amount of light scattered back to the
system resulting in higher signal-to-noise ratio (

#e1#

#s1#
#abstract2023.txt#


Time interleaved sigma-delta analog to digital converter seems to be a
potential solution for wide bandwidth analog to digital converter with
the lowest hardware complexity compared to other solutions using
parallel sigma-delta modulators. Its performance depends on the digital
filter and is very sensitive to the channel mismatch. This paper
summarizes our work on the digital signal processing for this kind of
converter, including filtering, decimation and channel mismatch
correction in order to reduce the implementation complexity while
minimizing the channel mismatch effect.

#e1#

#s1#
#abstract2024.txt#


We describe a high-speed software implementation of the eta /sub T/
pairing over binary supersingular curves at the 128-bit security level.
This implementation explores two types of parallelism found in modern
multi-core platforms: vector instructions and multiprocessing. We first
introduce novel techniques for implementing arithmetic in binary fields
with vector instructions. We then devise a new parallelization of
Miller's Algorithm to compute pairings. This parallelization provides
an algorithm for pairing computation without increasing storage costs
significantly. The combination of these acceleration techniques produce
serial timings at least 24% faster and parallel timings 66% faster than
the best previous result in an Intel Core platform, establishing a new
state-of-the art implementation of this pairing instantiation in this
platform.

#e1#

#s1#
#abstract2025.txt#


This paper presents a scalable core architecture based on a generic
systolic array. The size of this kind of cores can be adapted in
real-time to cover changing application requirements or to the
available area in a reconfigurable device. In this paper, the process
of scaling the core is performed by the replication of a single
processing element using run-time partial reconfiguration. Furthermore,
rather than restricting the proposed solution to a given application,
it is based on a generic systolic architecture which is adapted using a
design flow which is also proposed. The paper includes a related work
discussion, the proposal and definition of a systolic array
communication approach, which does not require the use of specific
macro structures and permits to achieve higher flexibility, and a
design flow used to adapt the generic architecture. Further, the paper
also includes an image filter application as a simple use case, along
with implementation results for Virtex 5 FPGA.

#e1#

#s1#
#abstract2026.txt#


As reconfigurable computing hardware and in particular FPGA-based
systems-on-chip comprise an increasing number of processor and
accelerator cores, supporting sharing and synchronization in a way that
is scalable and easy to program becomes a challenge. Transactional
memory (TM) is a potential solution to this problem, and an FPGA-based
system provides the opportunity to support TM in hardware (HTM).
Although there are many proposed approaches to HTM support for ASICs,
these do not necessarily map well to FPGAs. In particular in this work
we demonstrate that while signature-based conflict detection schemes
(essentially bit vectors) should intuitively be a good match to the
bit-parallelism of FPGAs, previous schemes result in either
unacceptable multicycle stalls, operating frequencies, or falseconflict
rates. Capitalizing on the reconfigurable nature of FPGA-based systems,
we propose an application-specific signature mechanism for HTM conflict
detection. Using both real and projected FPGA-based soft multiprocessor
systems that support HTM and implement threaded, shared-memory network
packet processing applications, relative to signatures with bit
selection we find that our application-specific approach (i) maintains
a reasonable operating frequency of 125 MHz, (ii) has an area overhead
of only 5%, and (iii) achieves a 9% to 71% increase in packet
throughput due to reduced false conflicts.

#e1#

#s1#
#abstract2027.txt#


We develop a generic framework for the analysis of programs with
recursive procedures and dynamic process creation. To this end we
combine the approach of weighted pushdown systems (WPDS) with the model
of dynamic pushdown networks (DPN). The resulting model, weighted
dynamic pushdown networks (WDPN), describes processes running in
parallel, each of them being able to perform pushdown actions, that may
spawn new processes as a side effect. As with WPDS, transitions are
labelled by weights to carry additional information. Starting from
techniques for WPDS and DPN, we derive a method to determine
meet-over-all-paths values for the paths between regular sets of
configurations of a WDPN. Using this method we are able to solve basic
data flow analysis problems in a parallel context.

#e1#

#s1#
#abstract2028.txt#


Stream processing applications such as algorithmic trading, MPEG
processing, and web content analysis are ubiquitous and essential to
business and entertainment. Language designers have developed numerous
domain-specific languages that are both tailored to the needs of their
applications, and optimized for performance on their particular target
platforms. Unfortunately, the goals of generality and performance are
frequently at odds, and prior work on the formal semantics of stream
processing languages does not capture the details necessary for
reasoning about implementations. This paper presents Brooklet, a core
calculus for stream processing that allows us to reason about how to
map languages to platforms and how to optimize stream programs. We
translate from three representative languages, CQL, Streamlt, and
Sawzall, to Brooklet, and show that the translations are correct. We
formalize three popular and vital optimizations, data-parallel
computation, operator fusion, and operator re-ordering, and show under
which conditions they are correct. Language designers can use Brooklet
to specify exactly how new features or languages behave. Language
implementors can use Brooklet to show exactly under which circumstances
new optimizations are correct. In ongoing work, we are developing an
intermediate language for streaming that is based on Brooklet. We are
implementing our intermediate language on System S, IBM's
high-performance streaming middleware.

#e1#

#s1#
#abstract2029.txt#


Modern software systems have frequently to face unexpected events,
reacting so to reach a consistent state. In the field of concurrent and
mobile systems (e.g., for web services) the problem is usually tackled
using long running transactions and compensations: activities
programmed to recover partial executions of long running transactions.
We compare the expressive power of different approaches to the
specification of those compensations. We consider (i) static recovery,
where the compensation is statically defined together with the
transaction, (ii) parallel recovery, where the compensation is
dynamically built as parallel composition of compensation elements and
(iii) general dynamic recovery, where more refined ways of composing
compensation elements are provided. We define an encoding of parallel
recovery into static recovery enjoying nice compositionality
properties, showing that the two approaches have the same expressive
power. We also show that no such encoding of general dynamic recovery
into static recovery is possible, i.e. general dynamic recovery is
strictly more expressive.

#e1#

#s1#
#abstract2030.txt#


Software development for Membrane Computing is growing up yielding new
applications. Nowadays, the efficiency of P systems simulators have
become a critical point when working with instances of large size. The
newest generation of GPUs (Graphics Processing Units) provide a
massively parallel framework to compute general purpose computations.
We present GPUs as an alternative to obtain better performance in the
simulation of P systems and we illustrate it by giving a solution to
the N-Queens problem as an example.

#e1#

#s1#
#abstract2031.txt#


This paper presents a program generator for fast software Viterbi
decoders for arbitrary convolutional codes. The input to the generator
is a specification of the code and a single-instruction multiple-data
(SIMD) vector length. The output is an optimized C implementation of
the decoder that uses explicit Intel SSE vector instructions. At the
heart of the generator is a small domain-specific language called VL to
express the structure of the forward pass. Vectorization is done by
rewriting VL expressions, which a compiler then translates into actual
code in addition to performing further optimizations specific to the
vector instruction set. Benchmarks show that the generated decoders
match the performance of available expert hand-tuned implementations,
while spanning the entire space of convolutional codes. An online
interface to the generator is provided at www.spiral.net.

#e1#

#s1#
#abstract2032.txt#


An implementation method of parallel finite element computation based
on overlapping domain decomposition was presented to improve the
parallel computing efficiency of finite element and lower the cost and
difficulty of parallel programming. By secondary processing the nodal
partition obtained by using Metis, the overlapping domain decomposition
of finite element mesh was gotten. Through the redundancy computation
of overlapping element, finite element governing equations could be
parallel formed independently. And the uniform distributed block
storage could be achieved conveniently. The interface to the DMSR data
format was developed to meet the need of Aztec parallel solution. And
the solver called the iterative solving subroutine of Aztec directly.
This implementation method reduced the change of the existed serial
program to a great extent. So the main frame of finite element
computation was kept. Tests show that this method can achieve high
parallel computing efficiency.

#e1#

#s1#
#abstract2033.txt#


A scalable parallel solver is developed to simulate the Earth's core
convection. With the help from the "multiphysics" data structure and
the restricted additive Schwarz preconditioning in PETSc the iterative
solution of the linear solver converges rapidly at every time-step. The
solver gains nearly 20 times speedup compared to a previous solver
using least-squares polynomial preconditioning in Aztec. We show the
efficiency and effectiveness of our new solver by giving numerical
results obtained on a BlueGene/L supercomputer with thousands of
processor cores.

#e1#

#s1#
#abstract2034.txt#


In this paper, we present an object-oriented concept of sparse matrix
and iterative linear solver for large scale parallel and sequential
finite element analysis of multi-field problems. With the present
concept, the partitioning and parallel solving of linear equation
systems can be easily realized, and the memory usage to solve coupled
multi-field problems is optimized. For the parallel computing, the
present objects are tailored to the domain decomposition approach for
both equation assembly and linear solver. With such approach, the
assembly of a global equation system is thoroughly avoided.
Parallelization is realized in the sparse matrix object by the means of
(1) enable the constructor of the sparse matrix class to use domain
decomposition data to establish the local domain sparse pattern and (2)
introduce MPI calls into the member function of matrix-vector
multiplication to collect local results and form global solutions. The
performance of these objects in C++ is demonstrated by a geotechnical
application of 3D thermal, hydraulic and mechanical (THM) coupled
problem in parallel manner.

#e1#

#s1#
#abstract2035.txt#


The bottleneck of most data analyzing systems, signal processing
systems, and intensive computing systems is matrix decomposition. The
Cholesky factorization of a sparse matrix is an important operation in
numerical algorithms field. This paper presents a Multi-phased Parallel
Cholesky Factorization (MPCF) algorithm, and then gives the
implementation on a multi-core machine. A performance result shows that
the system can reach 85.7 Gflop/s on a single PowerXCell processor and
bulk of computation can reach to 94% of peak performance.

#e1#

#s1#
#abstract2036.txt#


The large scale arbitrary cavity scattering problems can be tackled by
parallel computing approach. This paper proposes a new method to
distribute cavity scattering computing tasks on multiprocessors
computer based on the IPO and FMM. The samples of IPO and FMM pattern
can be divided into eight blocks, and this paper just analyzes the load
balancing of the block in the lower-left corner using the invariability
of angle. So this problem can be transformed into how to equally
distribute an N*N matrix on P processors. Two matrix partitioning
algorithms are introduced, and experimental results show that the
optimal sub-structure algorithm is able to balance the load among
multiprocessors effectively, thereby, improving the performance of the
entire system.

#e1#

#s1#
#abstract2037.txt#


Branch Prediction is a common function in nowadays microprocessor.
Branch predictor is duplicated into multiple copies in each core of a
multicore and many-core processor and makes prediction for multiple
concurrent running programs respectively. To evaluate the parallel
branch prediction in many-core processor, existed schemes generally use
a parallel simulator running in CPU which does not have a real passive
parallel running environment to support a many-core simulation and thus
has bad simulating performance. In this paper, we firstly try to use a
real many-core platform, GPU, to do a parallel branch prediction for
future general purpose many-core processor. We verify the new GPU based
parallel branch predictor against the traditional CPU based branch
predictor. Experiment result shows that GPU based parallel simulation
scheme is a promising way to faster simulating speed for future
many-core processor research.

#e1#

#s1#
#abstract2038.txt#


In order to realize the fast reconstruction processing of
all-reflective Fourier Transform Imaging Spectrometer (ARFTIS) data,
the authors implement reconstruction algorithms on GPU using Compute
Unified Device Architecture (CUDA). We use both CUDA 1DFFT library and
customization CUDA parallel kernel to accelerate the spectrum
reconstruction processing of ARFTIS. The results show that the CUDA can
drive hundreds of processing elements ('many-core' processors) of GPU
hardware and can enhance the efficiency of spectrum reconstruction
processing significantly. Farther, computer with CUDA graphic cards
will implement real-time data reconstruction of ARFTIS alone.

#e1#

#s1#
#abstract2039.txt#


The surface of the moon is scarred with millions of lunar crater, which
are the remains of collisions between an asteroid, comet, or meteorite
and the moon with different sizes and shapes. With the launch of
Chang's orbiter, it is available to detect the lunar crater at a high
resolution. However, the Chang's orbiter image is combined by different
path/orbit images. In this study, a batch processing scheme to detect
lunar crater rims from each path image is presented under a grid
environment. SGE (Sun Grid Engine) and OpenPBS (Open Portable Batch
System) are connected by Globus and MPICH-G2 as Linux PC Cluster
respectively. And the Globus GridFTP is used for parallel transfer of
different rows of Chang's orbiter images by MPICH-G2 model. The
detection algorithms on each node are executed respectively after the
parallel transfer. Thus, the lunar crater rims for the experimental
area are generated effectively.

#e1#

#s1#
#abstract2040.txt#


Preparing a CAD model for Finite Element (FE) analysis can be a
time-consuming task, where shape and mesh simplifications play an
important role. It is important that the simplified model has the same
mechanical properties as the original one, and that the deviation from
the original stays within a given tolerance. Most FE mesh
simplification algorithms are either fully or partially sequential, and
are therefore not suitable for architectures with high levels of
parallelism. Furthermore, the use of processors such as GPUs of IBMs
Cell BE require algorithms to be adapted to benefit from their
computational advantages. Here, we present an algorithm written for
parallel processors, and its implementation for the Cell BE.

#e1#

#s1#
#abstract2041.txt#


We explore three commodity parallel architectures: multi-core CPUs, the
Cell BE processor, and graphics processing units. We have implemented
four algorithms on these three architectures: solving the heat
equation, inpainting using the heat equation, computing the Mandelbrot
set, and MJPEG movie compression. We use these four algorithms to
exemplify the benefits and drawbacks of each parallel architecture.

#e1#

#s1#
#abstract2042.txt#


This work presents an approach for producing layouts of complex
workflows given in the Business Process Execution Language (BPEL). BPEL
is a verbose, hierarchical workflow language containing nested,
alternative and concurrent execution paths. This approach enhances the
Sugiyama algorithm by introducing special paths, which are constrained
to be drawn in parallel, and hence, orthogonally to the layers in the
Sugiyama model. To prove the feasibility of this approach, an extension
to the collaborative BPEL development system HOBBES was developed.
Collaboration enhances the need for visualizations of complex workflow
models, as team members have to coordinate their activities.

#e1#

#s1#
#abstract2043.txt#


A single-camera stereo vision system offers a greatly simplified
approach to image capture and analysis. In the original study by
Lovegrove and Brame, they proposed a novel optical system that uses a
single camera to capture two short, wide stereo images that can then be
analyzed for real-time obstacle detection. In this paper, further
analysis and refinement of the optical design results in two virtual
cameras with perfectly parallel central axes. A new prototype camera
design provides experimental verification of the analysis and also
provides insight into practical construction. The experimental device
showed that the virtual cameras' axes possessed a deviation from
parallel of less than 12 minutes (or 0.2 degrees ). The calculation of
distances to objects from the two overlapping images all showed errors
smaller than the pixel resolution limitation. In addition, a barrel
lens correction was used in processing the image to allow parallax
distance determination in the whole horizontal view of the images.

#e1#

#s1#
#abstract2044.txt#


Image reconstruction is one of the main challenges for fluorescence
tomography. For in vivo experiments on small animals, in particular,
the inhomogeneous optical properties and irregular surface of the
animal make free-space image reconstruction challenging because of the
difficulties in accurately modeling the forward problem and the finite
dynamic range of the photodetector. These two factors are fundamentally
limited by the currently available forward models and photonic
technologies. Nonetheless, both limitations can be significantly eased
using a signal processing approach. We have recently constructed a
free-space panoramic fluorescence diffuse optical tomography system to
take advantage of co-registered microCT data acquired from the same
animal. In this article, we present a data processing strategy that
adaptively selects the optical sampling points in the raw 2-D
fluorescent CCD images. Specifically, the general sampling area and
sampling density are initially specified to create a set of potential
sampling points sufficient to cover the region of interest. Based on
3-D anatomical information from the microCT and the fluorescent CCD
images, data points are excluded from the set when they are located in
an area where either the forward model is known to be problematic
(e.g., large wrinkles on the skin) or where the signal is unreliable
(e.g., saturated or low signal-to-noise ratio). Parallel Monte Carlo
software was implemented to compute the sensitivity function for image
reconstruction. Animal experiments were conducted on a mouse cadaver
with an artificial fluorescent inclusion. Compared to our previous
results using a finite element method, the newly developed parallel
Monte Carlo software and the adaptive sampling strategy produced
favorable reconstruction results.

#e1#

#s1#
#abstract2045.txt#


In this paper we propose a novel method for inferring an Inversion
Transduction Grammar (ITG) from a bilingual parallel corpus with
linguistic information from the source or target language. Our method
combines bilingual ITG parse trees with monolingual linguistic trees in
order to obtain a Syntax Augmented ITG (SAITG). The use of a modified
bilingual parsing algorithm with bracketing information makes possible
that each bilingual subtree has a correspondent subtree in the
monolingual parsing. In addition, several binarization techniques have
been tested for the resulting SAITG. In order to evaluate the effects
of the use of SAITGs in Machine Translation tasks, we have used them in
an ITG-based machine translation decoder. The results obtained using
SAITGs with the decoder for the IWSLT-08 Chinese-English machine
translation task produce significant improvements in BLEU.

#e1#

#s1#
#abstract2046.txt#


Discourse in instant messenger conversations (chats) with multiple
participants is often composed of several intertwining threads. Some
chat environments for Computer-Supported Collaborative Learning (CSCL)
support and encourage the existence of parallel threads by providing
explicit referencing facilities. The paper proposes a discourse model
for such chats, based on Mikhail Bakhtin's dialogic theory. It
considers that multiple voices (which do not limit to the participants)
inter-animate, sometimes in a polyphonic, counterpointal way. An
implemented system is also presented, which analyzes such chat logs for
detecting additional, implicit links among utterances and threads and,
more important for CSCL, for detecting the involvement
(inter-animation) of the participants in problem solving. The system
begins with a NLP pipe and concludes with inter-animation
identification in order to generate feedback and to propose grades for
the learners.

#e1#

#s1#
#abstract2048.txt#


A new method for correction of MRI motion artifacts induced by
corrupted /i k/-space data, acquired by multiple receiver coils such as
phased arrays, is presented. In our approach, a projections onto convex
sets (POCS)-based method for reconstruction of sensitivity encoded MRI
data (POCSENSE) is employed to identify corrupted /i k/-space samples.
After the erroneous data are discarded from the dataset, the
artifact-free images are restored from the remaining data using coil
sensitivity profiles. The error detection and data restoration are
based on informational redundancy of phased-array data and may be
applied to full and reduced datasets. An important advantage of the new
POCS-based method is that, in addition to multicoil data redundancy, it
can use a priori known properties about the imaged object for improved
MR image artifact correction. The use of such information was shown to
improve significantly /i k/-space error detection and image artifact
correction. The method was validated on data corrupted by simulated and
real motion such as head motion and pulsatile flow. c Wiley-Liss,
Inc.

#e1#

#s1#
#abstract2049.txt#


In this study, the sensitivity of the /i S//sub 2/-steady-state free
precession (SSFP) signal for functional MRI at 7 T was investigated. In
order to achieve the necessary temporal resolution, a three-dimensional
acquisition scheme with acceleration along two spatial axes was
employed. Activation maps based on /i S//sub 2/-steady-state free
precession data showed similar spatial localization of activation and
sensitivity as spin-echo echo-planar imaging (SE-EPI), but data can be
acquired with substantially lower power deposition. The functional
sensitivity estimated by the average z-values was not significantly
different for SE-EPI compared to the /i S//sub 2/-signal but was
slightly lower for the /i S//sub 2/-signal (6.74 +/- 0.32 for the TR =
15 ms protocol and 7.51 +/- 0.78 for the TR = 27 ms protocol) compared
to SE-EPI (7.49 +/- 1.44 and 8.05 +/- 1.67) using the same activated
voxels, respectively. The relative signal changes in these voxels upon
activation were slightly lower for SE-EPI (2.37% +/- 0.18%) compared to
the TR = 15 ms /i S//sub 2/-SSFP protocol (2.75% +/- 0.53%) and
significantly lower than the TR = 27 ms protocol (5.38% +/- 1.28%), in
line with simulations results. The large relative signal change for the
long TR SSFP protocol can be explained by contributions from multiple
coherence pathways and the low intrinsic intensity of the /i S//sub 2/
signal. In conclusion, whole-brain /i T//sub 2/-weighted functional MRI
with negligible image distortion at 7 T is feasible using the /i S//sub
2/-SSFP sequence and partially parallel imaging. Magn Reson Med
63:1015-1020, 2010. c Wiley-Liss, Inc.

#e1#

#s1#
#abstract2050.txt#


The field amplitude and frequency dependent complex ac susceptibility
chi (H/sub m/ , f) of four Y-Ba-Cu-O discs made by a top-seeded melt
growth technique has been measured at 77 K with the ac field applied
parallel to either the disc axis (the c axis of the grain) or its
radius (the ab plane of the grain). The results are interpreted by
considering the anisotropic nature of the critical-current density,
J/sub c/ . The axial susceptibility is dominated by J/sub c//sup ab,c/
for current within the ab plane and field parallel to the c axis,
whereas the radial susceptibility is dominated by J/sub c//sup c,ab/
for current along the c axis and field within the ab plane. It is
concluded that the ratio of J/sub c//sup ab,c/ to J/sub c//sup c,ab/ in
a thin, round surface layer ranges between 2 and 5 for discs of
different thicknesses, and both are much smaller than the value of
J/sub c//sup ab,ab/ associated with intrinsic pinning. In addition, the
measured susceptibilities of thicker grain discs suggest more granular
behavior, indicating the existence of microcracks within the sample
microstructure.

#e1#

#s1#
#abstract2051.txt#


This paper reports a systematic study of spatially extended atmospheric
plasma (SEAP) arrays employing many parallel plasma jets packed densely
and arranged in an honeycomb configuration. The work is motivated by
the challenge of using inherently small atmospheric plasmas to address
many large-scale processing applications including plasma medicine. The
first part of the study considers a capillary-ring electrode
configuration as the elemental jet with which to construct a 2D SEAP
array. It is shown that its plasma dynamics is characterized by strong
interaction between two plasmas initially generated near the two
electrodes. Its plume length increases considerably when the plasma
evolves into a high-current continuous mode from the usual bullet mode.
Its electron density is estimated to be at the order of 3.7 * 10/sup
12/ cm/sup -3/. The second part of the study considers 2D SEAP arrays
constructed from parallelization of identical capillary-ring plasma
jets with very high jet density of 0.47-0.6. Strong jet-jet
interactions of a 7-jet 2D array are found to depend on the excitation
frequency, and are effectively mitigated with the jet-array structure
that acts as an effective ballast. The impact range of the reaction
chemistry of the array exceeds considerably the cross-sectional
dimension of the array itself, and the physical reach of reactive
species generated by any single jet exceeds significantly the jet-jet
distance. As a result, the jet array can treat a large sample surface
without relative sample-array movement. A 37-channel SEAP array is used
to indicate the scalability with an impact range of up to 48.6 mm in
diameter, a step change in capability from previously reported SEAP
arrays. 2D SEAP arrays represent one of few current options as
large-scale low-temperature atmospheric plasma technologies with
distinct capability of directed delivery of reactive species and
effective control of the jet-jet and jet-sample interactions.

#e1#

#s1#
#abstract2052.txt#


In /sup 123/ I-IBZM brain SPECT, the main interest is the activity
uptake in the striatum relative to the background, and
semi-quantitative techniques using regions of interest are typically
used for this purpose. Uncertainties in the measured uptakes can
however be a problem due to low contrasts and high noise levels. Like
SPECT in general, IBZM SPECT should benefit from reconstruction methods
that include model-based compensation, but it is important that image
acquisition is optimized for this technique. An important factor is the
choice of collimator. In this study we compare four different
parallel-hole collimators for IBZM SPECT regarding overall quantitative
accuracy and measured uptake ratio as a function of image noise and
uncertainty. The collimators are low-energy high-resolution (LEHR),
low-energy general-purpose (LEGP), extended LEGP (ELEGP) and
medium-energy general-purpose (MEGP). The effect of three Butterworth
post-filters with cut-off frequencies of 0.3, 0.45 and 0.6 cm/sup -1/
(power factor 8) is also studied. All raw-data projections are produced
using Monte Carlo simulations. Of the investigated collimators, the one
that is most sensitive to the primary photons, ELEGP, proved to be the
most optimal for realistic noise levels. Butterworth post-filtering is
advantageous, and the cut-off frequency 0.45 cm/sup -1/ was the best
compromise in this study.

#e1#

#s1#
#abstract2053.txt#


In this paper, we proposed an architecture of embedded systems for
high-frame-rate real-time vision on the order of 1000 f/s, which
achieved both hardware reconfigurability and easy algorithm
implementation while fulfilling performance demands. The proposed
system consisted of an embedded microprocessor and field programmable
gate arrays (FPGAs). A coprocessor consisting of memory units, direct
memory access controller units, and image processing units were
implemented in each FPGA. While the number of units and functions are
reconfigurable by reprogramming the FPGAs, users can implement
algorithms without hardware knowledge. A descriptor method in which the
central processing unit gave instructions to each coprocessor through a
register array enabled task-level parallel processing as well as
pixel-level parallel processing in the processing units. The
specifications of an evaluation system developed based on the proposed
architecture, the results of performance evaluation, and application
examples using the system were shown.

#e1#

#s1#
#abstract2054.txt#


SU-8 cantilevers with a thickness of 2 mu m were fabricated using a dry
release method and two steps of SU-8 photolithography. The processing
of the thin SU-8 film defining the cantilevers was experimentally
optimized to achieve low initial bending due to residual stress
gradients. In parallel, the rotational deformation at the clamping
point allowed a qualitative assessment of the device release from the
fluorocarbon-coated substrate. The change of these parameters during
several months of storage at ambient temperature was investigated in
detail. The introduction of a long hard bake in an oven after
development of the thin SU-8 film resulted in reduced cantilever
bending due to removal of residual stress gradients. Further, improved
time-stability of the devices was achieved due to the enhanced
cross-linking of the polymer. A post-exposure bake at a temperature T
/sub PEB/ = 50 degrees C followed by a hard bake at T /sub HB/ = 90
degrees C proved to be optimal to ensure low cantilever bending and low
rotational deformation due to excellent device release and low change
of these properties with time. With the optimized process, the
reproducible fabrication of arrays with 2 mu m thick cantilevers with a
length of 500 mu m and an initial bending of less than 20 mu m was
possible. The theoretical spring constant of these cantilevers is k =
4.8 +/- 2.5 mN m/sup -1/ , which is comparable to the value for Si
cantilevers with identical dimensions and a thickness of 500 nm.

#e1#

#s1#
#abstract2055.txt#


The PC-based software programming used in complex or luxuriant image
processing algorithms is time consuming and resource wasting. As
appropriate processing for the image data indeed speedups complicated
algorithms, we focus on a crucial case - multilayered processes. In
this paper, we gauge deeply into the data flow of multilayered image
processing to avoid waiting for the result from every previous steps to
access the memory which occurs in many applicable algorithms. Based on
combining the parallel and pipelined properties to eliminate
unnecessary delays, we propose new visual pipeline architecture and use
field programmable gate array to implement our hardware scheme. For
verification, the multiscale Harris corner detector in cooperating with
shape context and thin-plate splines were combined to complete our
real-time experiment of the integrated hardware and software (H/S)
system for pattern recognition.

#e1#

#s1#
#abstract2056.txt#


Nuclear reaction gamma-ray diagnosis is one of the important techniques
used for studying confined fast-ions. The Joint European Torus (JET)
gamma-ray camera diagnostic provides information on the spatial
distribution of fast ions. The system is currently being upgraded and
should allow gamma-ray image measurements in high power deuterium JET
pulses, and eventually in deuterium-tritium discharges. In order to
fully exploit the diagnostic capabilities it is mandatory to develop a
reliable, maintainable, multi-channel spectroscopy data acquisition and
real-time processing (DAQP) system, which shares much of the common
development for other specific implementation like Gamma-ray
spectroscopy. The DAQP system is based on the Advanced
Telecommunications Computing Architecture (ATCA) and contains a 6
GFLOPS x86-based control unit and three transient recorder and
processing (TRP) modules, to cope with the two arrays of collimators
(10 horizontal + 9 vertical lines of sight), interconnected through PCI
Express (PCIe) links. Each TRP module features 8 channels of 13 bit
resolution sampling at 250 MHz, 4 GByte of local memory and two field
programmable gate arrays able to perform complex trigger managing modes
and allowing real time analyses (pulse height analyzer and pile-up
discrimination), minimizing data storage and transfer issues. The DAQP
system aims at overcoming the problem of storing large amount of data
during long discharges. A raw/processed mode is being developed where
the acquired raw data follows two parallel paths: besides being
directly stored in the on-board memory, it is processed and streamed in
real-time through PCIe links. This procedure is expected to greatly
reduce the amount of data and to give the possibility of allowing
continuous operation of the diagnostic. During commissioning and when
data validation is required, the 4 GB raw data will be executed on the
x86 control unit through a well known algorithm and the result cross
checked with the processed data.

#e1#

#s1#
#abstract2057.txt#


We have installed a serial-bus based fast readout control system in the
Belle experiment at the KEKB /i e//sup +//i e//sup -/ collider. Every
subdetector readout system at Belle comprises common readout boards
that initiate the readout sequence upon receiving a trigger signal for
every event. In order to control more than hundred of such readout
boards with a compact and scalable system, a cascadable tree structure
based on one-to-eight switch modules that are interconnected with
enhanced category-5 local area network cables was constructed. The
cable carries dedicated clock and trigger signals, and a pair of 10-bit
509 Mbps serial-bus lines to transfer trigger and control related
information. The 10-bit serial-bus is used as an 8-bit parallel lines
with 2-bit parity bits for most of the time, and when needed, it is
used to embed control sequences in other bit patterns to reset boards,
to collect status, and to measure the latency of the serial-bus. A
point-to-point path addressing method is implemented to access every
individual node at any location in the tree. This system has replaced
the old parallel and LEMO cable based fast control system for most of
the subdetector readout systems at Belle. Our experience on
installation and operation after running stably for more than two years
is reported.

#e1#

#s1#
#abstract2058.txt#


The High Level Trigger (HLT) and Data Acquisition System select about 2
kHz of events out of the 40 MHz of beam crossings. The selected events
are consolidated into files in onsite storage and then sent to
permanent storage for subsequent analysis on the Grid. For local and
full-chain tests a method to exercise the data-flow through the High
Level Trigger is needed in the absence of real data. In order to test
the system as much as possible under identical conditions as for
data-taking, the solution would be to inject data at the input of the
HLT at a minimum rate of 2 kHz. This is done via a software
implementation of the trigger system which sends data to the HLT. The
application has to simulate that the data it sends come from real LHCb
readout-boards. Data can come from several input streams, which are
selected according to probabilities or frequencies. Therefore the
emulator offers runs which are not only identical data-flows coming
from a sequence on tape, but physics-like pseudo-indeterministic
data-flow, including lumi events and candidate b-quark events. Both
simulation data and previously recorded real data can be re-played
through the system in this manner. As the data rate is high (100 MB/s),
care has been taken to optimize the emulator for throughput from the
Storage Area Network. The emulator can be run in stand-alone mode, but
even more interesting is that it can emulate any partition of LHCb in
parallel with the real hardware partition. In this mode it is fully
integrated into the standard run-control. The architecture,
implementation, and performance results of the emulator and full tests
will be presented. This emulator is a crucial part of the ongoing
data-challenges in LHCb. Results from these Full System Integration
Tests (FEST) will be presented, which helped to verify and benchmark
the entire LHCb data-flow.

#e1#

#s1#
#abstract2059.txt#


The LHCb experiment at CERN will have an Event Filter Farm (EFF)
composed of 2000 CPUs. These machines will form a pool of 50 sub-farms
with 30 to 40 nodes each, running a large amount of High Level Trigger
(HLT) tasks in parallel. Although these tasks are identical algorithms,
they can run at the same time being configured with different
parameters, such as run type (Physics, Cosmics, Test, etc.) or with
different subdetectors (partitions). The HLT is the second of the two
trigger levels in LHCb. Its selection algorithms reduce the incoming
data rate of 1 MHz to an output rate of 2 kHz. Selected events are sent
for mass storage and subsequent offline reconstruction and analysis.
These trigger processes running online are based on the same software
framework as the algorithms for offline analysis (Gaudi). The control
of the trigger farm was developed with an industrial SCADA system
(PVSS) which is used throughout the Experiment Control System (ECS).
The HLT algorithms are handled by the ECS like hardware devices, for
instance, high voltage channels. The integration of the HLT controls in
the overall ECS, which is modeled as finite state machines, will be
presented.

#e1#

#s1#
#abstract2060.txt#


Tubular bells are geometrically simple representatives of
three-dimensional vibrating structures. Under certain assumptions, a
tubular bell can be modeled as a rectangular plate with different types
of homogeneous boundary conditions. Suitable functional transformations
with respect to time and space turn the corresponding initial-boundary
value problem into a two-dimensional transfer function. An algorithmic
model follows according to the functional transformation method in
digital sound synthesis. As with simpler vibrating structures (strings,
membranes) the synthesis algorithms consist of a parallel arrangement
of second-order sections. Their coefficients are obtained by simple
analytic expressions directly from the physical parameters of the
tubular bell.

#e1#

#s1#
#abstract2061.txt#


A new rearrangeable nonblocking photonic multi-/i log//sub 2//i N/
network /i DM/(/i N/) is introduced. It is shown that /i DM/(/i N/)
network possesses many good properties simultaneously. These good
properties include all those of existing rearrangeable nonblocking
photonic multi-/i log//sub 2//i N/ networks and new ones such as /i
O/(log/i N/)-time fast parallel self-routing, nonblocking
multiple-multicast, and cost-effective crosstalk-free wavelength
dilation, which existing rearrangeable nonblocking multi-/i log//sub
2//i N/ networks do not have. The advantages of /i DM/(/i N/) over
existing multi-/i log//sub 2//i N/ networks, especially /i Log//sub
2//i (N/, 0, 2/sup lfloor log2 N/2 rfloor /), are achieved by employing
a two-level load balancing scheme-a combination of static load
balancing and dynamic load balancing. /i DM/(/i N/) and /i Log//sub
2//i (N/, 0, 2/sup lfloor log2 N/2 rfloor /) are about the same in
structure. The additional cost is for the intraplane routing
preprocessing circuits. Considering the extended capabilities of /i
DM/(/i N/) and current mature and cheap electronic technology, this
extra cost is well justified.

#e1#

#s1#
#abstract2063.txt#


Cardiac motion has been tracked using various methods, which vary in
their invasiveness and dimensionality. One such noninvasive modality
for cardiac motion tracking is ultrasound. Three-dimensional ultrasound
motion tracking has been demonstrated using detected data at low volume
rates. However, the effects of volume rate, kernel size, and data type
(raw and detected) have not been sufficiently explored. First
comparisons are made within the stated variables for 3-D speckle
tracking. Volumetric data were obtained in a raw, baseband format using
a matrix array attached to a high parallel receive beam count scanner.
The scanner was used to acquire phantom and human in vivo cardiac
volumetric data at 1000-Hz volume rates. Motion was tracked using
phase-sensitive normalized cross-correlation. Subsample estimation in
the lateral and elevational dimensions used the grid-slopes algorithm.
The effects of frame rate, kernel size, and data type on 3-D tracking
are shown. In general, the results show improvement of motion estimates
at volume rates up to 200 Hz, above which they become stable. However,
peak and pixel hopping continue to decrease at volume rates higher than
200 Hz. The tracking method and data show, qualitatively, good temporal
and spatial stability (for independent kernels) at high volume rates.

#e1#

#s1#
#abstract2064.txt#


Verbal communication between individuals requires the parallel
evolution of a vocal system capable of emitting different sounds and of
an auditory system able to recognize each vocal pattern. In this work
we present the evolution of a population of twins where the selection
pressure is based on the ability of learning a communication pattern
which allows verbal transmission of information. The fitness of each
pair of twins (i.e. individuals having the same genotype) is based on
the percentage of correct recognition of the perceived sounds. Results
indicate the evolved communication system, in absence of noise, rapidly
evolves and reaches almost 100% correct classifications, while, even in
presence of a strong noise either in the channel, or in the sound
generation parameters, the system can obtain a very good performance
(approximately 80% correct classifications in the worst case).

#e1#

#s1#
#abstract2065.txt#


Appraising the Renewable Energy Sources (RES) options' contribution to
Sustainable Development (SD) is a complex task, considering the
different aspects of SD the imprecision and uncertainty of the related
information as well as the qualitative aspects embodied, that cannot be
represented by numerical values. The objective of the paper is to show
how energy policy objectives towards SD and RES options are related and
assessed using linguistic variables. The presented method extends the
numerical multicriteria method TOPSIS for processing linguistic
information, eliminating in parallel the loss of information caused by
the approximation procedures and the ambiguity of the fuzzy ranking
method's selection. Moreover, the application of the method's
computerized software to the RES interventions proposed in the Hellenic
National Action Plan for Greenhouse Gases is presented and discussed.
[All rights reserved Elsevier].

#e1#

#s1#
#abstract2066.txt#


Purpose: P. R. Edholm, R. M. Lewitt, and B. Lindholm, "Novel properties
of the Fourier decomposition of the sinogram," in Proceedings of the
International Workshop on Physics and Engineering of Computerized
Multidimensional Imaging and Processing [Proc. SPIE 671, 8-18 (1986)]
described properties of a parallel beam projection sinogram with
respect to its radial and angular frequencies. The purpose is to
perform a similar derivation to arrive at corresponding properties of a
fan-beam projection sinogram for both the equal-angle and equal-spaced
detector sampling scenarios. Methods: One of the derived properties is
an approximately zero-energy region in the two-dimensional Fourier
transform of the full fan-beam sinogram. This region is in the form of
a double-wedge, similar to the parallel beam case, but different in
that it is asymmetric with respect to the frequency axes. The authors
characterize this region for a point object and validate the derived
properties in both a simulation and a head CT data set. The authors
apply these results in an application using algebraic reconstruction.
Results: In the equal-angle case, the domain of the zero region is
(q,k) for which |k/(k-q)|#R/L, where q and k are the frequency
variables associated with the detector and view angular positions,
respectively, R is the radial support of the object, and L is the
source-to-isocenter distance. A filter was designed to retain only
sinogram frequencies corresponding to a specified radial support. The
filtered sinogram was used to reconstruct the same radial support of
the head CT data. As an example application of this concept, the
double-wedge filter was used to computationally improve region of
interest iterative reconstruction. Conclusions: Interesting properties
of the fan-beam sinogram exist and may be exploited in some
applications.

#e1#

#s1#
#abstract2067.txt#


Purpose: A rotating multi-segment slant-hole (RMSSH) collimator is able
to provide much higher (~3 times for four-segment collimator with 30
degrees slant angle) sensitivity than a parallel-hole (PH) collimator
with the similar spatial resolution for imaging small organs such as
the heart and the breast. In this article, the authors evaluated the
performance of myocardial perfusion SPECT (MPS) using a RMSSH
collimator compared to MPS using the low-energy high-resolution
parallel-hole collimators. Methods: The authors conducted computer
simulation studies using the NURBS-based cardiac-torso phantom,
receiver operative characteristic (ROC) analysis using the channelized
Hotelling observer, physical phantom experiments, and pilot patient
studies to evaluate the performance of MPS using a rotating
four-segment slant-hole (R4SSH) collimator with respect to MPS using a
PH collimator. Results: In the simulation study, the R4SSH MPS provides
images with superior contrast-noise trade-off than those of PH MPS with
the same acquisition time. The defect detectability in terms of the
largest area under the ROC curve for R4SSH MPS is significantly higher
than those of PH MPS with p-values #0.01. In the phantom experiments,
the R4SSH MPS images with 7.5 min acquisition had similar noise level
and overall image quality as those of PH MPS with 21 min acquisition.
Pilot patient studies showed that with the same acquisition time, the
R4SSH SPECT using a single-head camera gave images with similar quality
as those of PH SPECT using a dual-head camera. Conclusions: The RMSSH
SPECT has a potential to improve the coronary artery disease detection
and workflow of SPECT imaging acquisition due to the high sensitivity
property of the RMSSH collimator.

#e1#

#s1#
#abstract2068.txt#


This paper presents a fast split-radix- (2×2)/(8×8)
algorithm for computing the 2-D discrete Hartley transform (DHT) of
length /i N/ */i N/ with /i N/ = /i q/ middot 2 /i m/, where /i q/ is
an odd integer. The proposed algorithm decomposes an /i N/ * /i N/ DHT
into one /i N/ /2 * /i N/ /2 DHT and 48 /i N/ /8 * /i N/ /8 DHTs. It
achieves an efficient reduction on the number of arithmetic operations,
data transfers and twiddle factors compared to the
split-radix-(2*2)/(4*4) algorithm. Moreover, the characteristic of
expression in simple matrices leads to an easy implementation of the
algorithm. If implementing the above two algorithms with fully parallel
structure in hardware, it seems that the proposed algorithm can
decrease the area complexity compared to the split-radix-(2*2)/(4*4)
algorithm, but requires a little more time complexity. An application
of the proposed algorithm to 2-D medical image compression is also
provided.

#e1#

#s1#
#abstract2072.txt#


This paper presents the results of a case study which investigates the
use of an embedded soft-core processor to perform Built-In Self-Test
(BIST) of the logic resources in Xilinx Virtex-5 Field Programmable
Gate Arrays (FPGAs). We show that the approach reduces the complexity
of an external BIST controller and the number of external
reconfigurations, making it particularly appealing for in-system
testing of high-reliability and fault-tolerant systems with FPGAs.
However, the overall test time is not improved due to an increase in
the size of the required configuration files as a consequence of the
inclusion of the soft-core embedded processor logic, whose relative
irregularity results in less effective compression of configuration
data files.

#e1#

#s1#
#abstract2076.txt#


We report on direct inscription of type-II waveguides in bulk
titanium-doped sapphire with an ultrafast chirpedpulse oscillator.
Ti/sup 3+/:Sapphire is of particular interest due to its large emission
bandwidth which enables a broadband tunability and generation of
ultra-short pulses. However, its lasing threshold is high and powerful
high brightness pump sources are required. The fabrication of a
waveguide in Ti/sup 3+/:Sapphire could thus enable the fabrication of
low-threshold tunable lasers and broadband fluorescence sources. The
latter are of interest for optical coherence tomography where the
obtainable resolution scales with the bandwidth of the light source.
The fabricated waveguides are formed in-between two laser induced
damage regions. This technique has been applied to other crystalline
materials (e.g. LiNbO/sub 3/) but not in Ti/sup 3+/:Sapphire, yet. The
size of the structural changed regions is strongly dependent on the
writing laser polarization. These damage regions of changed structure
cause a stress-field inside the crystalline lattice which consequently
increases the refractive index to form a waveguide. The written
structures exhibit a strong birefringence and two waveguides that
support orthogonal polarized modes are formed between each pair of
damage lines. Linearly polarized light parallel to the crystal's
surface is guided between the two damage regions while a waveguide for
the orthogonal polarization is formed underneath. The propagation
properties of the waveguides are characterized by their near-field
profiles and insertion losses with respect to the writing parameters.
Further the fluorescence output power is measured and the emission
spectra of the waveguides are compared to the bulk material.

#e1#

#s1#
#abstract2077.txt#


The development of a real-time optical waveform measurement technique
with quantum-limited sensitivity, unlimited record lengths and an
instantaneous bandwidth scalable to terahertz frequencies would be
beneficial in the investigation of many ultrafast optical phenomena.
Currently, full-field (amplitude and phase) optical measurements with a
bandwidth greater than 100 GHz require repetitive signals to facilitate
equivalent-time sampling methods or are single-shot in nature with
limited time records. Here, we demonstrate a bandwidth- and time-record
scalable measurement that performs parallel coherent detection on
spectral slices of arbitrary optical waveforms in the 1.55 mu m
telecommunications band. External balanced photodetection and
high-speed digitizers record the in-phase and quadrature-phase
components of each demodulated spectral slice, and digital signal
processing reconstructs the signal waveform. The approach is passive,
extendable to other regions of the optical spectrum, and can be
implemented as a single silicon photonic integrated circuit.

#e1#

#s1#
#abstract2098.txt#


The radio-frequency (rf)-to-microwave impedance spectra of solution
grown ZnO nanorods have been measured from 0.1 to 50 GHz using vector
network analysis. To increase interaction with rf/microwave fields, the
nanorods were assembled by dielectrophoresis into arrays on coplanar
waveguides. The average complex impedance frequency response per
nanorod in an array was accurately modeled as a simple three-element
circuit composed of the inherent nanorod resistance in series with a
parallel resistor-capacitor representing the contact. The nanorod
resistance dominates at high frequencies while the contact impedance
dominates at low frequencies, permitting a quantitative separation of
contact effects from nanorod properties. The average inherent
resistivity of a nanorod was found to be ~10/sup -2/ Omega cm,
indicating the nanorods were unintentionally highly doped. Accuracy of
the inherent resistance measurement was limited by the highly
conductive nature of the nanorods used and the upper limit of the
experimental frequency range. Determination of the nanorod resistance
becomes more accurate for higher resistivity nanorods, so high
frequency impedance spectroscopy will provide an increasingly valuable
electrical characterization technique as the ability to synthesize more
intrinsic (i.e., lower unintentional dopant density) ZnO nanorods
improves.

#e1#

#s1#
#abstract2099.txt#


The need for computational resources capable of processing geospatial
data has accelerated the uptake of geospatial web services. Several
academic and commercial organizations now offer geospatial web services
for data provision, coordinate transformation, geocoding and several
other tasks. These web services adopt specifications developed by the
Open Geospatial Consortium (OGC) - the leading standardization body for
Geographic Information Systems. In parallel with efforts of the OGC,
the Grid computing community has published specifications for
developing Grid applications. The Open Grid Forum (OGF) is the main
body that promotes interoperability between Grid computing systems.
This study examines the integration of Grid services and geospatial web
services into workflows for Geoscientific processing. An architecture
is proposed that bridges web services based on the abstract geospatial
architecture (ISO19119) and the Open Grid Services Architecture (OGSA).
The paper presents a workflow management system, called SAW-GEO, that
supports orchestration of Grid-enabled geospatial web services. An
implementation of SAW-GEO is presented, based on both the Simple
Conceptual Unified Flow Language (SCUFL) and the Business Process
Execution Language for Web Services (WS-BPEL or BPEL for short).

#e1#

#s1#
#abstract2101.txt#


The paper deals with design and performance analysis of algorithms that
utilize parallel signal-processing methods and SIMD technology for
multiply-and-add algorithm for digital audio signal processing. This
algorithm is used for summing the gained input signals on output buses
in applications for distributing, mixing, effect-processing, and
switching multi-format digital audio signal in an audio signal network
on desktop processors platforms. The subjective evaluation of latency
caused by principle of the real-time digital audio processing is also
studied in the paper Results of an analysis of speed-up and real-time
performance of several summing algorithms are presented in the paper as
well as subjective evaluation of the latency depending on the audio
buffer size.

#e1#

#s1#
#abstract2102.txt#


The aim of this paper is compare the effect of using different
topologies or connections between separate colonies in island based
parallel implementations of the Ant Colony Optimization applied to the
Minimum Weight Vertex Cover Problem. We investigated the sequential Ant
Colony Optimization algorithms applied to the Minimum Weight Vertex
Cover Problem before. Parallelization of population based algorithms
using the island model is of great importance because it often gives
super linear increase in performance. We observe the behavior of
different parallel algorithms corresponding to several topologies and
communication rules like fully connected, replace worst, ring and
independent parallel runs. We also propose a variation of the algorithm
corresponding to the ring topology that maintains the diversity of the
search, but still moves to areas with better solutions and gives
slightly better results even on a single processor with threads.

#e1#

#s1#
#abstract2103.txt#


The methodology for reshaping and enlarging the Brillouin gain spectral
response is proposed in this article as a technique to develop
all-fiber optical active devices. It is based on the superposition of
the Brillouin scattering spectra from several optical fibers, which are
connected in serial, parallel, or both. Additionally, the overall
Brillouin gain spectral response can be tailored or customized by
tuning the temperature or the strain on the fiber. To experimentally
demonstrate the method, we propose a fiber device with a tailored like
"W" spectral response. The bandwidth is 190 MHz at 3 dB, centered in
10.765 GHz from the pump wavelength. c Wiley Periodicals, Inc.

#e1#

#s1#
#abstract2104.txt#


We demonstrated a real-time display of processed OCT images using a
linear-in-wavenumber (linear-k) spectrometer and a graphics processing
unit (GPU). We used the linear-k spectrometer with optimal combination
of a diffractive grating with 1200 lines/mm and a F2 equilateral prism
in the 840 nm spectral region, to avoid calculating the re-sampling
process. The calculations of the FFT (fast Fourier transform) were
accelerated by the low cost GPU with many stream processors, which
realized highly parallel processing. A display rate of 27.9 frames per
second for processed images (2048 FFT size * 1000 lateral A-scans) was
achieved in our OCT system using a line scan CCD camera operated at
27.9 kHz.

#e1#

#s1#
#abstract2105.txt#


We report on the feasibility of rapid, high resolution, 3-dimensional
swept source optical coherence tomography (3D SSOCT) to detect early
airway injury changes following smoke inhalation exposure in a rabbit
model. The SSOCT system obtains 3-D helical scanning using a
microelectromechanical system (MEMS) motor based endoscope. Real-time
2-D data processing and image display at the speed of 20 frames per
second are achieved by adopting the technique of shared-memory parallel
computing. Longitudinal images are reconstructed via an image
processing algorithm to remove motion artifacts caused by ventilation
and pulse. We demonstrate the ability of the SSOCT system to detect
increases in tracheal and bronchial airway thickness that occurs
shortly after smoke exposure.

#e1#

#s1#
#abstract2106.txt#


In this work we modified light illumination of the laser optoacoustic
(OA) imaging system to improve the 3D visualization of human forearm
vasculature. The computer modeling demonstrated that the new
illumination design that features laser beams converging on the surface
of the skin in the imaging plane of the probe provides superior OA
images in comparison to the images generated by the illumination with
parallel laser beams. We also developed the procedure for vein/artery
differentiation based on OA imaging with 690 nm and 1080 nm laser
wavelengths. The procedure includes statistical analysis of the
intensities of OA images of the neighboring blood vessels. Analysis of
the OA images generated by computer simulation of a human forearm
illuminated at 690 nm and 1080 nm resulted in successful
differentiation of veins and arteries. In vivo scanning of a human
forearm provided high contrast 3D OA image of a forearm skin and a
superficial blood vessel. The blood vessel image contrast was further
enhanced after it was automatically traced using the developed
software. The software also allowed evaluation of the effective blood
vessel diameter at each step of the scan. We propose that the developed
3D OA imaging system can be used during preoperative mapping of forearm
vessels that is essential for hemodialysis treatment.

#e1#

#s1#
#abstract2107.txt#


We present results from a clinical case study on imaging breast cancer
using a real-time interleaved two laser optoacoustic imaging system
co-registered with ultrasound. The present version of Laser
Optoacoustic Ultrasonic Imaging System (LOUIS) utilizes a commercial
linear ultrasonic transducer array, which has been modified to include
two parallel rectangular optical bundles, to operate in both ultrasonic
(US) and optoacoustic (OA) modes. In OA mode, the images from two
optical wavelengths (755 nm and 1064 nm) that provide opposite
contrasts for optical absorption of oxygenated vs deoxygenated blood
can be displayed simultaneously at a maximum rate of 20 Hz. The
real-time aspect of the system permits probe manipulations that can
assist in the detection of the lesion. The results show the ability of
LOUIS to co-register regions of high absorption seen in OA images with
US images collected at the same location with the dual modality probe.
The dual wavelength results demonstrate that LOUIS can potentially
provide breast cancer diagnostics based on different intensities of OA
images of the lesion obtained at 755 nm and 1064 nm. We also present
new data processing based on deconvolution of the LOUIS impulse
response that helps recover original optoacoustic pressure profiles.
Finally, we demonstrate the image analysis tool that provides automatic
detection of the tumor boundary and quantitative metrics of the
optoacoustic image quality. Using a blood vessel phantom submerged in a
tissue-like milky background solution we show that the image contrast
is minimally affected by the phantom distance from the LOUIS probe
until about 60-65 mm. We suggest using the image contrast for
quantitative assessment of an OA image of a breast lesion, as a part of
the breast cancer diagnostics procedure.

#e1#

#s1#
#abstract2108.txt#


The data acquisition speed in photoacoustic computed tomography (PACT)
is limited by the laser repetition rate and the number of parallel
ultrasound detecting channels. Reconstructing PACT image with a less
number of measurements can effectively accelerate the data acquisition
and reduce the system cost. Recently emerged Compressed Sensing (CS)
theory enables us to reconstruct a compressible image with a small
number of projections. This paper adopts the CS theory for
reconstruction in PACT. The idea is implemented as a non-linear
conjugate gradient descent algorithm and tested with phantom and in
vivo experiments.

#e1#

#s1#
#abstract2109.txt#


Explicit congestion control schemes use router feedback to overcome
limitations of the standard mechanisms of the Transmission Control
Protocol (TCP). These approaches require additional packet processing
in every router and therefore raise the question whether, and how, this
can be achieved in high-speed routers. This paper investigates the
realization complexity of these router functions of two such schemes,
the TCP Quick-Start extension and the Explicit Control Protocol (XCP).
Our focus lies on the implementation using a network processor. We show
that synchronization issues among parallel processing entities have to
be considered, and that this affects the router performance. We develop
and compare different synchronization mechanisms for highly parallel
packet processing. Our prototype implementation on an Intel IXP network
processor allows to quantify the impact on throughput and delay caused
by the additional packet processing in the fast path. The measurements
reveal that Quick-Start and XCP processing is feasible at multiple
Gbit/s line speed, with Quick-Start being simpler to scale. We expect
similar results for the implementation of the Rate Control Protocol
(RCP), which is another router-assisted congestion control scheme,
requiring no elaborate synchronization. Finally, we study the
implementation using programmable logic and show the applicability of
XCP and in particular Quick-Start even at significantly higher line
speeds.

#e1#

#s1#
#abstract2110.txt#


Multistage Interconnection Networks (MINs) are used to interconnect
different processing modules in various parallel systems or on high
bandwidth networks. In this paper an integrated performance methodology
is presented. A new approximate performance model for self-routing MINs
consisting of symmetrical switches which are subject to a backpressure
blocking mechanism is analyzed. Based on this, the steady-state
distribution of the queue utilization is estimated and then all
important performance metrics are calculated. Moreover, a general
evaluation factor which helps in choosing a better performance MIN in
comparison with other similar MIN architecture specifications is
defined. The model was exemplified for the case of symmetrical single-
and double-buffered MINs. It provides accurate results and converges
very quickly. The obtained results were validated by extensive
simulations and were compared to existing related work in the
literature.

#e1#

#s1#
#abstract2111.txt#


We address the probabilistic generalization of weighted flow time on
parallel machines. We present some results for situations which ask for
"long-term robust" schedules of n jobs (tasks) on m parallel machines
(processors): on any given day, only a random subset of jobs needs to
be processed. The goal is to design robust a priori schedules (before
we know which jobs need to be processed) which, on a long-term horizon,
are optimal (or near optimal) with respect to total weighted flow time.
The originality of this work is that probabilities are explicitly
associated with data such that further classical properties of a task
(processing time and weight) we consider a probability of presence.
After motivating this investigation we analyze the computational
complexity, analytical properties, and solution procedures for these
problems. Special care is also devoted to assess experimentally the
performance of a priori strategies. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2112.txt#


A multicast routing infrastructure is proposed as a core feature of
SpiNNaker, a massively parallel computer for the real-time simulation
of large-scale spiking neural networks. The infrastructure is
implemented using a communications router, based on an event-driven
routing scheme, on each multicore processing node in the system. The
design considerations emphasize the difference between the requirements
of neural network communications and those of conventional computer
networks and on-chip networks. The focus of the design is on neural
modelling flexibility, power-efficiency, fault-tolerance and the
communication throughput of the router.

#e1#

#s1#
#abstract2113.txt#


Precise control of diffraction peaks of a hologram is indispensable in
holographic femtosecond laser processing. To obtain the uniform
diffraction peaks, an adaptive optimization due to the diffraction
peaks measured by an image sensor was proposed. It used a one-photon
absorption. However, the structure processed by a femtosecond laser
pulse was based on multi-photon absorptions. Therefore, a mismatch
between the optimized diffraction peaks and the processed structures
was observed. An adaptive optimization method using second harmonics
induced by parallel pulse irradiations to a nonlinear optical crystal
is proposed to solve this mismatch.

#e1#

#s1#
#abstract2114.txt#


A parallel processing of two-photon polymerization structuring is
demonstrated with spatial light modulator. Spatial light modulator
generates multi-focus spots on the sample surface via phase modulation
technique controlled by computer generated hologram pattern. Each focus
spot can be individually controlled in position and laser intensity
with computer generated hologram pattern displayed on spatial light
modulator. The multi-focus spots two-photon polymerization achieves the
fabrication of asymmetric structure. Moreover, smooth sine curved
polymerized line with amplitude of 5 mu m and a period of 200 mu m was
obtained by fast switching of CGH pattern.

#e1#

#s1#
#abstract2115.txt#


Parallel femtosecond laser processing using a computer-generated
hologram (CGH) displayed on a spatial light modulator, called
holographic femtosecond laser processing, has advantages of high
throughput and high light-use efficiency of the laser pulse energy. We
demonstrate two types of the holographic femtosecond laser processing
that are implemented with a Fourier transform CGH and a Fresnel
transform CGH. We present some demonstrations of two-dimensional and
three-dimensional parallel laser processing.

#e1#

#s1#
#abstract2116.txt#


High intensity intensity ultraviolet (UV) and vacuum ultraviolet (VUV)
radiation provide a singular dominant narrow-band emission at various
wavelengths( lambda ) between 108 - 351 nm. The use of
dielectric-barrier discharges in its embodiment of an excimer lamp as a
photon-source provides a novel method to induce surface modification.
From its in relatively humble beginnings in ozone generation, the
excimer lamp has found new applications in the field of low-temperature
processing of surfaces. Herein, a 15 year perspective of work done at
the Materials & Devices Group at University College London between 1992
and 2007 is presented. The excimer lamps' application to the
modification of surfaces for materials processing include:
photo-induced formation of high- kappa dielectric thin films and more
recently the UV-induced photo-doping of silicon substrates, amongst
others. With its robust yet inexpensive setup and flexibility of
geometric configurations, they are easily coupled in parallel resulting
in the provision of high photon fluxes over large areas. These sources
also have an incoherent and almost monochromatic selectivity for
application to process chemical pathway specific tasks by simple
variation of the discharge gas mixture. These sources are an
interesting addition to and an alternative to lasers for scalable
industrial applications and have potential for a myriad of applications
across different fields.

#e1#

#s1#
#abstract2117.txt#


This letter proposes a fast parallel architecture and redundancy
reduction algorithm for H.264/AVC intra4*4 prediction to speed up intra
frame coding. A significant reduction in execution time is achieved
without losing video quality. Only 204 cycles are required to process a
macroblock (MB). Compared with the dedicated intra prediction,
processing speed is enhanced by 79%.

#e1#

#s1#
#abstract2118.txt#


Hydrous ruthenium oxide/carbon black nanocomposites were prepared by
impregnation of the carbon blacks by differently aged inorganic RuO/sub
2/ sols, i.e. of different particle size. Commercial Black Pearls
2000/sup reg / (BP) and Vulcan/sup reg / XC-72 R (XC) carbon blacks
were used. Capacitive properties of BP/RuO/sub 2/ and XC/RuO/sub 2/
composites were investigated by cyclic voltammetry (CV) and
electrochemical impedance spectroscopy (EIS) in H/sub 2/SO/sub 4/
solution. Capacitance values and capacitance distribution through the
composite porous layer were found different if high- (BP) and low- (XC)
surface-area carbons are used as supports. The aging time (particle
size) of Ru oxide sol as well as the concentration of the oxide solid
phase in the impregnating medium influenced the capacitive performance
of prepared composites. While the capacitance of BP-supported oxide
decreases with the aging time, the capacitive ability of XC-supported
oxide is promoted with increasing oxide particle size. The increase in
concentration of the oxide solid phase in the impregnating medium
caused an improvement of charging/discharging characteristics due to
pronounced pseudocapacitance contribution of the increasing amount of
inserted oxide. The effects of these variables in the impregnation
process on the energy storage capabilities of prepared nanocomposites
are envisaged as a result of intrinsic way of population of the pores
of carbon material by hydrous Ru oxide particle. [All rights reserved
Elsevier].

#e1#

#s1#
#abstract2119.txt#


Deep Space Optical Communications (DSOC) impose challenging
requirements on detector sensitivity and bandwidth. The current
state-of-the art of high-repetition rate, high-power lasers recommends
using near-infrared (NIR) 1064nm wavelengths for specific DSOC tasks.
Large photonic arrays with integrated beam acquisition, tracking and/or
communication capabilities, and smart pixel architecture should allow
the implementation of more reliable and robust DSOC systems.
Integration of smart pixel technology for parallel data read,
acquisition and processing is currently available in silicon. Therefore
it would be desirable to monolithic ally integrate the photodetectors
with the electronics. However, silicon has a weak absorption at 1064nm.
One elegant approach to increase its absorption efficiency is to trap
the photons inside the silicon using the cavity resonance effect
(resonant cavity enhancement or RCE). We present in this paper the
challenges of developing resonant cavity single-photon detector arrays
for applications to DSOC. The metrics of the main process parameters to
fabricate resonant cavity detectors is analyzed and critical process
steps are developed and evaluated. We conclude that such detector
arrays are feasible using current state-of-the-art CMOS technology,
provided that suitable process control protocols are developed. We
report a 10X performance enhancement at NIR wavelengths for the first
generation of resonant cavity single-photon detector prototypes, less
than 150ps timing performance in photon-starved mode and 20-30ps for
multi-photon hits.

#e1#

#s1#
#abstract2120.txt#


Image-based holographic stereogram rendering methods for holographic
video have the attractive properties of moderate computational cost and
correct handling of occlusions and translucent objects. These methods
are also subject to the criticism that (like other stereograms) they do
not present accommodation cues consistent with vergence cues and thus
do not make use of one of the significant potential advantages of
holographic displays. We present an algorithm for the Diffraction
Specific Coherent Panoramagram -- a multi-view holographic stereogram
with correct accommodation cues, smooth motion parallax, and visually
defined centers of parallax. The algorithm is designed to take
advantage of parallel and vector processing in off-the-shelf graphics
cards using OpenGL with Cg vertex and fragment shaders. We introduce
wavefront elements - "wafels" - as a progression of picture element
"pixels", directional element "direls", and holographic element
"hogels". Wafel apertures emit controllable intensities of light in
controllable directions with controllable centers of curvature,
providing accommodation cues in addition to disparity and parallax
cues. Based on simultaneously captured scene depth information, sets of
directed variable wavefronts are created using nonlinear chirps, which
allow coherent diffraction of the beam across multiple wafels. We
describe an implementation of this algorithm using a commodity graphics
card for interactive display on our Mark II holographic video display.

#e1#

#s1#
#abstract2121.txt#


Nanocrystalline cuboidal ceria has been synthesized by low-temperature
hydrothermal reaction of cerium nitrate hexahydrate with hexamethylene
tetramine. The particles have been doped with La and Gd by adding
aqueous solution of the nitrate salts of the metals to the reaction
mixture. The pure and doped particles are cubic in crystal structure
and 10-25 nm in size. The pure and La-doped ceria are cuboidal in
morphology, whereas the Gd-doped particles are irregular in shape.
High-resolution TEM imaging and image simulation indicates that atomic
level steps are present on the particle surfaces. The particles are
faceted parallel to the {1 1 1} and {1 0 0} crystallographic planes and
a continuous switching takes place between the two possible surface
facets. It appears that the surface energies of the {1 1 1} and {1 0 0}
facets are quite similar in magnitude and the interplay of surface
energy determines the particle shape. Chemically sensitive imaging and
spectroscopy shows that the dopants are homogeneously distributed
within the particles and that the oxidation state of Ce is a mixture of
+3 and +4. No preferential segregation either of the dopant or the
oxidation state was observed. However, since the facet switching does
depend on the chemistry of the dopant, there must be an affect on the
atomic scale. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2124.txt#


Background: ChlP-Seq, which combines chromatin immunoprecipitation
(ChIP) with high-throughput massively parallel sequencing, is
increasingly being used for identification of protein-DNA interactions
in vivo in the genome. However, to maximize the effectiveness of data
analysis of such sequences requires the development of new algorithms
that are able to accurately predict DNA-protein binding sites. Results:
Here, we present SIPeS (Site Identification from Paired-end
Sequencing), a novel algorithm for precise identification of binding
sites from short reads generated by paired-end solexa ChlP-Seq
technology. In this paper we used ChlP-Seq data from the Arabidopsis
basic helix-loop-helix transcription factor ABORTED MICROSPORES (AMS),
which is expressed within the anther during pollen development, the
results show that SIPeS has better resolution for binding site
identification compared to two existing ChlP-Seq peak detection
algorithms, Cisgenome and MACS. Conclusions: When compared to Cisgenome
and MACS, SIPeS shows better resolution for binding site discovery.
Moreover, SIPeS is designed to calculate the mappable genome length
accurately with the fragment length based on the paired-end reads.
Dynamic baselines are also employed to effectively discriminate closely
adjacent binding sites, for effective binding sites discovery, which is
of particular value when working with high-density genomes.

#e1#

#s1#
#abstract2125.txt#


Recently, a new approximation algorithm for the nonpreemptive
scheduling of independent jobs on m identical parallel processors has
appeared in the literature. The algorithm, named MPS (multiprocessor
scheduling), combines partial solutions which satisfy suitable
properties. Its performance ratio is bounded by z+1/z - 1/mz, where z
represents the number of initial partial solutions provided by the
algorithm. This note presents an advanced estimate of z and,
consequently, an improved worst-case performance ratio of the MPS
algorithm.

#e1#

#s1#
#abstract2126.txt#


Mg-doped GaN nanowires have been successfully grown on Si(111)
substrates by magnetron sputtering through ammoniating Ga/sub 2/O/sub
3//Mg thin films at 900 degrees C for 15 min. The growth of the GaN
nanowires was investigated as a function of ammoniating time so as to
study the influence of ammoniating time on the structural properties of
GaN samples in particular by X-ray diffraction (XRD), X-ray
photoelectron spectroscopy (XPS), FT-IR spectrophotometer, scanning
electron microscopy (SEM), high-resolution transmission electron
microscopy (TEM), and photoluminescence (PL) spectrum. The results
demonstrate that ammoniating time has great influence on the
microstructure, morphology and optical properties of GaN nanowires. GaN
nanowires ammoniated at 900 degrees C for 15 min are straight and
smooth with uniform thickness along spindle direction and high
crystalline quality, 50 nm in diameter and 20 mu m in length with good
emission properties, and the growth direction of the nanowire is
parallel to [100] orientation. A clear blue-shift of the band-gap
emission has occurred due to Mg doping. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2127.txt#


Background: The selection of genes that discriminate disease classes
from microarray data is widely used for the identification of
diagnostic biomarkers. Although various gene selection methods are
currently available and some of them have shown excellent performance,
no single method can retain the best performance for all types of
microarray datasets. It is desirable to use a comparative approach to
find the best gene selection result after rigorous test of different
methodological strategies for a given microarray dataset. Results: FiGS
is a web-based workbench that automatically compares various gene
selection procedures and provides the optimal gene selection result for
an input microarray dataset. FiGS builds up diverse gene selection
procedures by aligning different feature selection techniques and
classifiers. In addition to the highly reputed techniques, FiGS
diversifies the gene selection procedures by incorporating gene
clustering options in the feature selection step and different data
pre-processing options in classifier training step. All candidate gene
selection procedures are evaluated by the .632+ bootstrap errors and
listed with their classification accuracies and selected gene sets.
FiGS runs on parallelized computing nodes that capacitate heavy
computations. FiGS is freely accessible at
http://gexp.kaist.ac.kr/figs. Conclusion: FiGS is an web-based
application that automates an extensive search for the optimized gene
selection analysis for a microarray dataset in a parallel computing
environment. FiGS will provide both an efficient and comprehensive
means of acquiring optimal gene sets that discriminate disease states
from microarray datasets.

#e1#

#s1#
#abstract2128.txt#


We study the problem of scheduling jobs with release times and
due-dates on a single machine with the objective to minimize the
maximal job lateness. This problem is strongly NP-hard, however it is
known to be polynomially solvable for the case when the processing
times of some jobs are restricted to either p or 2p, for some integer
p. We present a polynomial-time algorithm when job processing times are
less restricted; in particular, when they are mutually divisible. Our
algorithm works for the case when for any pair of jobs, if one is
longer than another then the due-date of the former job is no larger
than that of the latter one.

#e1#

#s1#
#abstract2129.txt#


The recent advances in computer animation, motion image processing,
robotics and so on prompted us to analyze computational complexity of
four-dimensional pattern processing. Thus, the research of
four-dimensional automata as a computational model of four-dimensional
pattern processing has also been meaningful. From this viewpoint, we
introduced a four-dimensional alternating Turing machine (4-ATM)
operating in parallel. In this paper, we continue the investigations
about 4-ATM's, deal with a four-dimensional synchronized alternating
Turing machine (4-SATM), and investigate some properties of 4-SATM's
which each sidelength of each input tape is equivalent. The main topics
of this paper are: (1) hierarchies based on the number of processes of
4-SATM's, and (2) recognizability of connected pictures by 4-SATM's.

#e1#

#s1#
#abstract2130.txt#


The DAANNS (diffusion algorithm asynchronous nearest neighborhood)
proposed in this paper is an asynchronous algorithm which finds the
unbalanced nodes in a network of processors automatically, and makes
them balanced in a way that load difference among the processors is 0.5
units. In this paper the DAANNS algorithm has been evaluated by
comparing it with RID (receiver initiated diffusion) algorithm across a
range of network topologies including ring, hypercube, and torus where
the no of nodes have been varied from 8 to 128. All the experiments
were performed on Intel parallel compiler. After simulation we have
noticed that the DAANNS performed very well as compared to RID.

#e1#

#s1#
#abstract2131.txt#


In the framework of heavy mid-level processing for high speed imaging,
a nonlinear bi-dimensional network is proposed, allowing the
implementation of active curve algorithms. Usually this efficient type
of algorithm is prohibitive for real-time image processing due to its
calculus charge and the inadequate structure for the use of serial or
parallel architectures. Another kind of implementation philosophy is
proposed here, by considering the active curve generated by a
propagation phenomenon inspired from biological modeling. A
programmable nonlinear reaction-diffusion system is proposed under
front control and technological constraints. Geometric multiscale
processing is presented and this opens a discussion about electronic
implementation. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2133.txt#


We describe a new iteration of the StereoJet process, which has been
simplified by changes in materials and improved by the conversion from
linear to circular polarization. A prototype StereoJet process for
producing full color stereoscopic images, described several years ago
by Scarpetti et al., was developed at the Rowland Institute for
Science, now part of Harvard University. The system was based on the
inkjet application of inks comprising dichroic dyes to polaroid
vectograph sheet, a concept explored earlier by Walworth and Chiulli at
the Polaroid Research Laboratories. Vectograph sheet comprised two
oppositely oriented layers of stretched polyvinyl alcohol (PVA)
laminated to opposite surfaces of a cellulose triacetate support sheet.
The two PVA layers were oriented at +45 and -45 degrees, respectively,
with respect to the running edge of the support sheet. A left-eye and
right-eye stereoscopic image pair were printed sequentially on the
respective surfaces, and the resulting stereoscopic image viewed with
conventional linearly polarized glasses having +45 and -45 degree
orientation. StereoJet, Inc. has developed new, simplified technology
based on the use of PVA substrate of the type used in sheet polarizer
manufacture with orientation parallel to the running edge of the
support. Left- and right-eye images are printed at 0 and 90 degrees,
then laminated in register. Addition of a thin layer of 1/4-wave
retarder to the front surface converts the image pair's respective
orientations to right- and left-circular polarization. The full color
stereoscopic images are viewed with circularly polarized glasses.

#e1#

#s1#
#abstract2134.txt#


Introduction: Many operative specialties rely on the use of microscopes
and endoscopes for visualizing operative fields in minimally invasive
or microsurgical interventions. Conventional optical devices present
relevant details only to the surgeon or one assisting professional.
Advances in information technology have propagated stereoscopic visual
information to all operative theatre personnel in real time, which, in
turn, adds to the load of complex technical devices to be handled and
maintained in operative theatres. Material and methods: In the last six
years, we have been using conventional (SD, 720 * 576 pixels) and high
definition (1280 * 720 pixels) stereoscopic video cameras attached to
conventional operative microscopes either in parallel to direct
visualization of the operative field or as all-digital processing of
the operative image for the surgeon including all other staff. Aspects
included the type of display used, image quality, time delay due to
image processing, visual comfort, time consumption in set-up, ease of
use as well as robustness of the system. Results: General acceptance of
stereoscopic display technology is high as all staff members are able
to share the same visual information. The use of stereo cameras in
parallel to direct visualization or as only visualization device mostly
depended on image quality and personal preference of the surgeon.
Predominant general factors are robustness, ease of use and additional
time consumption imposed by setup and handling. Visual comfort was
noted as moderately important as there was wide variability between
staff members. Type of display used and post-processing issues were
regarded less important. Time delay induced by the video chain was
negligible. Conclusion: The additional information given by
stereoscopic video processing in real time outweighs the extra effort
for handling and maintenance. However, further integration with
existing technology and with the general workflow enhances acceptance
especially in units with high turnover of operative procedures.

#e1#

#s1#
#abstract2135.txt#


Rear-projected screens such as those in Digital Light Projection (DLP)
televisions suffer from an image quality problem called hot spotting,
where the image is brightest at a point dependent on the viewing angle.
In rear-projected multi-screen configurations such as the StarCAVE at
Calit2, this causes discontinuities in brightness at the edges where
screens meet, and thus in the 3D image perceived by the user. In the
StarCAVE we know the viewer's position in 3D space and we have
programmable graphics hardware, so we can mitigate this effect by
performing post-processing in the inverse of the pattern, yielding a
homogenous image at the output. Our implementation improves brightness
homogeneity by a factor of 4 while decreasing frame rate by only 1-3
fps.

#e1#

#s1#
#abstract2136.txt#


As a video coding standard, H.264 achieves high compress rate while
keeping good fidelity. But it requires more intensive computation than
before to get such high coding performance. A hierarchical multi-level
parallelisms (HMLP) framework for H.264 encoder is proposed which
integrates four level parallelisms - frame-level, slice-level,
macroblock-level and data-level into one implementation. Each level
parallelism is designed in a hierarchical parallel framework and mapped
onto the multi-cores and SIMD units on multi-core architecture.
According to the analysis of coding performance on each level
parallelism, we propose a method to combine different parallel levels
to attain a good compromise between high speedup and low bit-rate. The
experimental results show that for CIF format video, our method
achieves the speedup of 33.57x-42.3x with 1.04x-1.08x bitrate
increasing on 8-core Intel Xeon processor with SIMD technology.

#e1#

#s1#
#abstract2138.txt#


This paper describes a new system for sharing a 3D space on workbenches
placed at different locations. It consists of flatbed-type
autostereoscopic displays based on the one-dimensional (horizontal
parallax only) integral imaging (1D-II) method and multi-viewpoint
video cameras. Possible applications of the system include a tool for
remote instruction, where an instructor can show how to assemble a
product from given parts to a worker who is not in front of the
instructor but in a different room or factory. The idea of sharing the
3D space at different locations is not new. In the previous
applications such as mixed reality, however, since the depth of the
reconstructed 3D space was not clearly restricted, it is difficult to
improve the image quality of the reconstructed space. In the
application presented in this paper, we can obtain a reasonable level
of the image quality because the depth of the reconstructed space is
clearly limited by the size of the parts on the 3D workbench. A new
multi-viewpoint video camera was designed for the application. In this
design, each image sensor was placed parallel to the lens with a shift
in each direction of the XY coordinate system in a horizontal plane and
only a limited region of each image was used to reconstruct the 3D
space. An experimental system called the "3D hand-area space sharing
system" was implemented using integral imaging 3D displays as well as
multi-viewpoint video cameras. As a result, we ascertained that larger
display, namely, 8K4K, and interpolation technology are necessary for a
multi-viewpoint video camera for the 3D hand-area space sharing system.

#e1#

#s1#
#abstract2139.txt#


Sequential Monte Carlo (SMC) simulations are widely used to solve
problems associated with complex probability distribution. Intensive
computations are their main drawbacks, which restrict to be applied to
real time applications, and thus efficient parallelism under high
performance computing environment is crucial to effective
implementations, especially for intelligent computer vision systems.
The combination of auxiliary variables importance sampling with Markov
chain Monte Carlo (MCMC) resampling for pipelining data are proposed in
this paper so as to minimize executive time, whilst improve the
estimation accuracy. Experimental resultion a network of workstations
composed of simple off-the-shelf hardware components show that the
hybrid parallel scheme provides a bottleneck free to reduce executive
time with increasing particles, compared to the conventional SMC and
MCMC based parallel schemes.

#e1#

#s1#
#abstract2140.txt#


This paper presents an approach to reconstruct solid models from
triangular meshes of STL files. First, suitable slicing planes should
be selected for extracting parallel intersection contours, which will
be used for solid model reconstruction. Usually, a suitable flat region
of triangular meshes of the STL model is selected as the bottom
surface, and it can be fitted into a plane from the selected flat
region. The flat region is separated by a mesh segmentation method,
which uses a specified small threshold dihedral angle to divide all
triangular facets into separated regions. Next, a series of parallel
slicing contours are obtained by cutting the STL model through
specified parallel cutting planes. Slicing contours are originally
composed of a lot of line segments, which should be simplified and
refitted into 2D NURBS curves for data reduction and contour smoothing.
The number of points on each slicing contour is reduced by comparing
the variation of included angles of each two adjacent line segments.
Reduced points of each slicing contour are fitted into a NURBS curve in
commercial CAD software. Finally, with a series of parallel 2D NURBS
curves, the solid model of the STL facets is established by loft
operations supplied in almost all popular CAD software. The established
solid model can be used for other post processing such as finite
element mesh generation.

#e1#

#s1#
#abstract2141.txt#


We present a novel parallel algorithm for fast continuous collision
detection (CCD) between deformable models using multi-core processors.
We use a hierarchical representation to accelerate these queries and
present an incremental algorithm that exploits temporal coherence
between successive frames. Our formulation distributes the computation
among multiple cores by using fine-grained front-based decomposition.
We also present efficient techniques to reduce the number of elementary
tests and analyze the scalability of our approach. We have implemented
the parallel algorithm on eight core and 16 core PCs, and observe up to
7* and 13* speedups respectively, on complex benchmarks. [All rights
reserved Elsevier].

#e1#

#s1#
#abstract2142.txt#


Image annotation can be formulated as a classification problem.
Recently, Adaboost learning with feature selection has been used for
creating an accurate ensemble classifier. We propose dynamic Adaboost
learning with feature selection based on parallel genetic algorithm for
image annotation in MPEG-7 standard. In each iteration of Adaboost
learning, genetic algorithm (GA) is used to dynamically generate and
optimize a set of feature subsets on which the weak classifiers are
constructed, so that an ensemble member is selected. We investigate two
methods of GA feature selection: a binary-coded chromosome GA feature
selection method used to perform optimal feature subset selection, and
a bi-coded chromosome GA feature selection method used to perform
optimal-weighted feature subset selection, i.e. simultaneously perform
optimal feature subset selection and corresponding optimal weight
subset selection. To improve the computational efficiency of our
approach, master-slave GA, a parallel program of GA, is implemented.
k-nearest neighbor classifier is used as the base classifier. The
experiments are performed over 2000 classified Corel images to validate
the performance of the approaches. [All rights reserved Elsevier].

#e1#

#s1#
#abstract2147.txt#


Accurate measurement of spatially variant noise in MR images acquired
using parallel imaging techniques is challenging. Image-based noise
measurement methods such as the subtraction method proposed by the
National Electrical Manufacturers Association or the multiple
acquisition method often cannot be applied in vivo due to motion and/or
dynamic contrast changes. Based on the Karhunen-Loeve transform and
random matrix theory, we propose a novel method to accurately assess
the noise variance in image series bearing temporal redundancy. The
method fits the probability density function of eigenvalues from the
temporal covariance matrix of the image series to the Marcenko-Pastur
distribution. The accuracy of our method was validated using numerical
simulation and an MR noise measurement experiment. The ability of this
method to derive the /i g/-factor map of a static phantom was validated
against the multiple acquisition method. The method was applied to in
vivo cardiac and brain image series and the results agreed with
subtraction and multiple acquisition methods, respectively. This new
image-based noise measurement method provides a practical means of
retrospectively evaluating the noise level and/or /i g/-factor map from
multiframe image series. Magn Reson Med 63:782-789, c
Wiley-Liss, Inc.

#e1#

#s1#
#abstract2148.txt#


High-resolution (~0.22 mm) images are preferably acquired on whole-body
7T scanners to visualize minianatomic structures in human brain. They
usually need long acquisition time (~12 min) in three-dimensional
scans, even with both parallel imaging and partial Fourier samplings.
The combined use of both fast imaging techniques, however, leads to
occasionally visible undersampling artifacts. Spiral imaging has an
advantage in acquisition efficiency over rectangular sampling, but its
implementations are limited due to image blurring caused by a strong
off-resonance effect at 7T. This study proposes a solution for
minimizing image blurring while keeping spiral efficient. Image
blurring at 7T was, first, quantitatively investigated using computer
simulations and point-spread functions. A combined use of multishot
spirals and ultrashort echo time acquisitions was then employed to
minimize off-resonance-induced image blurring. Experiments on phantoms
and healthy subjects were performed on a whole-body 7T scanner to show
the performance of the proposed method. The three-dimensional brain
images of human subjects were obtained at echo time = 1.18 ms,
resolution = 0.22 mm (field of view = 220 mm, matrix size = 1024), and
in-plane spiral shots = 128, using a home-developed ultrashort echo
time sequence (acquisition-weighted stack of spirals). The total
acquisition time for 60 partitions at pulse repetition time = 100 ms
was 12.8 min without use of parallel imaging and partial Fourier
sampling. The blurring in these spiral images was minimized to a level
comparable to that in gradient-echo images with rectangular
acquisitions, while the spiral acquisition efficiency was maintained at
eight. These images showed that spiral imaging at 7T was feasible. c
Wiley-Liss, Inc.

#e1#

#s1#
#abstract2149.txt#


Carbon nanotubes (CNTs) irradiated by Ar ion beams at elevated
temperature were studied. The irradiation-induced defects in CNTs are
greatly reduced by elevated temperature. Moreover, the two types of CNT
junctions, the crossing junction and the parallel junction, were
formed. And the CNT networks may be fabricated by the two types of CNT
junctions. The formation process and the corresponding mechanism of CNT
networks are discussed. [All rights reserved Elsevier].

#e1#